# GEO Toolbox — full content > Generative engine optimization (GEO) measured across eight AI engines. Track AI search visibility on ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews, Google AI Mode, Bing Copilot, and Grok. The index version of this file is at https://geotoolbox.ai/llms.txt. # Blog ## Ahrefs Alternatives (2026): The Honest, Job-Based List > The best Ahrefs alternatives for 2026, by job: broadest all-in-one, keyword research, backlinks, content, and AI-visibility tracking. Prices checked August 2026. - Canonical: https://geotoolbox.ai/blog/ahrefs-alternatives - Published: 2026-08-18 · Updated: 2026-08-22 If you are looking for an Ahrefs alternative, you have probably hit one of three walls: the price and the jump between tiers, the credits that run out mid-project, or the fact that there is no real trial to test it first. The answer most listicles dodge is that there is no single drop-in replacement. Ahrefs bundles several jobs into one login, and the right alternative depends on which of those jobs you actually do. So this guide is organized by job, not by a ranking that happens to put the highest-paying affiliate first. Every price here was checked in August 2026, we flag where we have a commercial interest, and we tell you where a free tool beats a paid one. Find the row that matches your problem and skip the rest. ## Why People Look for Ahrefs Alternatives Ahrefs is one of the most respected SEO tools on the market, and almost nobody leaves it because the data is bad. Its backlink index is the one the rest of the field gets measured against. People leave for a few specific reasons, and knowing which one is yours tells you which alternative to pick. **The price, and the jump between tiers.** The Starter plan is $29 a month, but it is a limited, spot-check tier, billed monthly only. Real research starts at the Lite plan at $129 a month, and the next step up, Standard, is $249. For a solo consultant or a small site, that is a big leap for a tool you may open a few times a week. **The credits, which is the complaint most lists skip.** This is one of the loudest themes in the forums, louder than most feature gaps. Ahrefs meters report credits, crawl credits, and API units behind the scenes, and a deep research session burns through them fast. One user described running out of credits twelve days into a billing cycle during a client deadline, then being stuck until the reset. When you spend more time watching a usage meter than doing the work, the tool is the wrong shape for the job. **No real free trial.** Ahrefs does not offer a classic time-boxed trial. There is Ahrefs Free (the renamed Webmaster Tools), but the tools you actually want, Site Explorer and Site Audit, only work on sites you have verified you own. You cannot run a competitor's domain through it before you pay, which is exactly what most people want a trial for. There is also a simpler reason: you only report on your own site. If you are checking your own rankings and traffic rather than running competitive research at volume, Ahrefs is more tool than you need. Search Console and a free keyword tool cover most of what you actually open it for. None of these means Ahrefs is bad. They mean it is a poor fit for a specific job, and there is almost always a cheaper or more focused option for that job. ## Why There's No One-to-One Ahrefs Replacement Ahrefs is not one product. It is backlink analysis, keyword research, rank tracking, site audits, content tools, and now an AI-visibility add-on, bundled into one login. No competitor matches all of that in the same shape, which is why "what's the best Ahrefs alternative?" has no single answer. Which one wins comes down to the job you are hiring it for. That also means switching has a cost, and the honest lists say so. People who cancel and rebuild their stack from three cheaper tools often come back, and the reasons are consistent: the backlink index is smaller and slower to refresh, the keyword database is thinner, and each tool has its own learning curve. One tool that does several jobs well can beat three tools that each do one, once you count the hours of stitching them together. So the useful question is not "what replaces Ahrefs?" It is "which one thing am I actually paying Ahrefs for, and what does that one thing cost somewhere else?" The rest of this guide is organized around that question, one job at a time. ## Ahrefs Alternatives at a Glance

Almost every single-job alternative undercuts the suite. The last column is the one the classic SEO suites treat as an afterthought.

The reasoning behind each pick is below, and a fuller pricing table with the caveats is near the end. ## The Best Ahrefs Alternatives, by the Job You're Hiring For ### If You Want the Broadest All-in-One Suite The closest all-around rival to Ahrefs is [Semrush](/go/semrush?ref=ahrefs-alternatives-semrush), and it is the tool most people mean when they say "Ahrefs but broader." Ask the answer engines yourself: when we ran the question "what's the best alternative to Ahrefs?" through ChatGPT, Gemini, Perplexity, and Claude, Semrush was the one name every engine named, and the top pick for three of the four (Claude led with SE Ranking, then Semrush). Google's own AI Overview for the query names Semrush as the best direct competitor. The reason is scope. Semrush matches Ahrefs on keyword research, rank tracking, backlinks, and audits, then adds the things Ahrefs is thin on: deep PPC and competitor-ad intelligence, local SEO, social, and a more built-out content toolkit. Its keyword database is broad, and in our own testing it cataloged far more of a young site's early rankings than Ahrefs did. If you are consolidating several tools into one login, this is the one that covers the most jobs. The SEO plan is $139 a month (less on annual billing, from about $117.33), with a [7-day free trial](/go/semrush-seo?ref=ahrefs-alternatives-semrush-trial) so you can run your own accounts through it before committing. We put the two suites head to head on our own sites in [Semrush vs Ahrefs](https://geotoolbox.ai/blog/semrush-vs-ahrefs) if you are choosing between them specifically. **Watch for:** Semrush is single-seat, so extra users are a paid add-on, and cancellation runs through an emailed confirmation link that catches people out, so calendar your trial end date. Its AI-visibility layer is also a separate purchase (more on that below). Our full [Semrush review](https://geotoolbox.ai/blog/semrush-review) covers the billing mechanics before you start a trial. **Our pick for most people leaving Ahrefs for something broader:** [Semrush](/go/semrush-seo?ref=ahrefs-alternatives-verdict) is the closest all-in-one match and the name that came up across every engine we tested. Run the 7-day trial on your busiest account and decide on your own data, then calendar the end date so the billing does not decide for you. ### If You Want an All-in-One for Less If you want most of the suite at a lower price, [SE Ranking](/go/seranking?ref=ahrefs-alternatives-seranking) is the value pick. It covers keyword research, rank tracking, site audits, competitor analysis, and backlink monitoring, and the Core plan is $129 a month ($103 billed annually). The 14-day trial does not ask for a card. Report generation is in-plan, but white-label client reporting is a separate add-on, the Agency Pack, from $69/mo billed annually. One thing most listicles get wrong: SE Ranking includes basic AI-search visibility in its plans, but the fuller AI Search Toolkit is a paid add-on (about $79 a month), not a bundled freebie. On AI tracking it is not meaningfully cheaper than Semrush's add-on, so switch for the lower price on the core suite and the no-card trial, not for free AI tracking. **Watch for:** the backlink index is smaller and less fresh than Ahrefs, so if links are your main job, this is not the pick. SE Ranking's price also scales with how often you want rankings refreshed, so a high-frequency setup costs more than the sticker. ### If You Just Need Keyword Research For keyword work without the suite, [Mangools](/go/mangools?ref=ahrefs-alternatives-mangools) is the long-standing favorite. Its KWFinder tool is genuinely pleasant to use, the difficulty scores hold up against real-world ranking, and the Basic plan is $29 a month billed annually, roughly a fifth of Ahrefs Lite. It is now sold as an AI plus SEO bundle that includes AI Search Watcher, which tracks brand mentions across ChatGPT, Claude, Gemini, Grok, and more. There is no free trial, but there is a free limited account and a 48-hour money-back guarantee. **Watch for:** it is a keyword and SERP tool, not a full platform. The backlink and audit tooling is shallow, so treat it as a focused instrument. Neil Patel's Ubersuggest is the other cheap option here, at $29 a month or a one-time $290, and it is the better fit if you want basic audits bundled with keyword ideas. ### If Backlinks Are the Whole Job If links are your entire job, the best Ahrefs alternative is often to stay on Ahrefs. Its backlink index is widely regarded as the freshest and highest-quality of the major tools, and it is still the benchmark most link data gets compared against. (Semrush now advertises a larger raw link count, so this is about quality and refresh rate, not the biggest number.) Trading down to save money on link research usually means missing links, especially newer ones. If you do need a cheaper link-only tool, Majestic is the specialist, built around its Trust Flow and Citation Flow metrics, from about $49.99 a month. And if you want a broad suite whose backlink index is the least of a step down, that is Semrush again. Treat the budget link tools as a supplement to Search Console, not a full replacement for a deep index. ### If You Only Report on Your Own Site For watching your own rankings and traffic, you may not need a paid tool at all. Google Search Console shows the queries you actually rank for and the clicks you actually get, Google Analytics covers behavior, Keyword Planner gives volume ranges, and Looker Studio ties it into a report. That stack is free and, for your own site, more truthful than any third-party estimate (it is your first-party data, though it has its own sampling and attribution limits). Paid suites earn their price on competitive research, the data about sites you do not own. If you are not doing that at volume, you are renting a race car to drive to the shop. Our [best free SEO tools](https://geotoolbox.ai/blog/best-free-seo-tools) guide maps the free stack job by job. Come back to a paid tool when competitor research, not self-reporting, becomes the work. ### If Content Optimization Is the Job If what you actually open Ahrefs for is planning and grading content, a dedicated content tool does it better. [Surfer](/go/surfer?ref=ahrefs-alternatives-surfer) scores a draft against the pages already ranking and gives concrete term and structure targets, at around $99 a month for its standard plan. Clearscope is the pricier, more editorial-grade option that in-house content teams tend to prefer, and it has added AI-cited-pages tracking to connect published work to how LLMs answer. **Watch for:** these are on-page optimization tools, not research suites. They tell you how to shape a page, not which pages to build. Pair one with a keyword tool. Our [content optimization tools](https://geotoolbox.ai/blog/best-content-optimization-tools) roundup compares the field. ### If You Live in Technical Audits For technical SEO, Screaming Frog is the standard desktop crawler and it is cheap for what it does: $279 a year for the full version, free for crawls under 500 URLs. It goes deeper on on-page and crawl issues than any suite's site audit, at the cost of a steeper learning curve. It is not a full Ahrefs replacement, but as an alternative to Ahrefs' Site Audit specifically, it is the tool practitioners reach for. ### If You Need Competitor PPC Data Ahrefs is thin on paid-search intelligence, and if that is the job, SpyFu is the budget specialist. It surfaces the keywords a rival buys and ranks for, their ad history, and estimated spend, from about $39 a month. For a solo marketer who needs competitor ad data more than a full suite, it is a fraction of the cost, though its data skews toward larger, US-heavy advertisers. Semrush is the deeper option here if PPC intelligence is a core, ongoing need rather than an occasional lookup. ### The Job the SEO Suites Bury: AI and GEO Visibility Here is the disclosure up front, because it matters for this section: we build one of these tools. geotoolbox is our AI-visibility product, so read this as an interested party, not a neutral referee. The reason it belongs on an Ahrefs-alternatives list is that AI-search visibility is the job the classic suites treat as a bolt-on. Ahrefs answers it with Brand Radar, sold as two standalone plans: Select Platforms, from $199 a month per platform (you add platforms individually), or All Platforms at $699 a month for the full set. Semrush charges $99 a month per domain for its AI Visibility Toolkit unless you buy the $199-a-month Semrush One bundle. SE Ranking includes a basic version in its plans and sells a fuller toolkit as an add-on; Mangools and Ubersuggest bundle AI-search tracking into their plans. The point is not that one is cheaper. It is that if the thing you want to know is whether ChatGPT, Perplexity, Google's AI Overviews, and the rest mention and cite you, you are usually paying for a whole SEO suite to reach one module. A tool built for that job does two things a bolt-on rarely does well. It tracks mention and citation rates across the answer engines your buyers actually use, and, the part most suites skip, it checks whether those engines can even fetch your pages. A blocked AI crawler quietly sinks your visibility no matter how good the content is, and few, if any, of the suite add-ons test for that. [geotoolbox](https://geotoolbox.ai/pricing) starts at $99 a month ($79 annual, 7-day trial) and reaches up to eight engines on higher tiers, including Claude, Copilot, and Grok, which the suite add-ons did not all cover as of mid-2026. You can check the reachability half for free before paying for anything: our [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) tells you in about a minute whether those engines can actually crawl your pages, no account needed. Our [AI visibility tools roundup](https://geotoolbox.ai/blog/best-ai-visibility-tools) compares the whole field, including us, with the same disclosure made here.
![geotoolbox free AI-readiness scan result showing a grade-A score with passing checks: robots.txt valid, AI bot discoverability with 0 AI crawlers blocked, and a valid sitemap.](/blog/ahrefs-alternatives/geotoolbox-ai-readiness.png)
The reachability half, free: our AI-readiness scan checks whether the answer engines' crawlers are actually allowed to fetch your pages before you pay to track visibility.
**Watch for:** geotoolbox does one job. It is not an Ahrefs alternative for backlinks, keyword research, or audits, and it is a newer, smaller tool than the suites. And treat every AI-visibility number from any vendor, ours included, with some skepticism. These tools sample nondeterministic engines, so a share-of-voice figure is an estimate with a margin, not a meter reading. Buy on whether a tool measures the engines you care about and checks crawler access, not on a single headline percentage. If you are unsure the category is worth a line item yet, read [what AI visibility actually is](https://geotoolbox.ai/blog/what-is-ai-visibility) first. ## When Ahrefs Is Still the Right Call Not every reason to shop around survives contact with the alternatives. If backlink research is the core of your work, Ahrefs is often still the right tool, and this guide would be dishonest to pretend otherwise. Its link index is fresher and higher-quality than anything on this list, its keyword and rank data are strong, and the interface is cleaner than most, which is why it earns the loyalty it does. The trade is this: you leave Ahrefs for a broader marketing suite (Semrush), a cheaper all-in-one (SE Ranking), a focused keyword tool (Mangools), or a dedicated AI-visibility layer (a tool like ours). You do not leave it because its core data is weak. If you find yourself missing links or second-guessing a budget tool's index a month after switching, that is the signal that links were your real job all along, and Ahrefs was doing it. ## A Pricing Reality Check Two cautions before the table. First, these prices move constantly and often differ by region and billing period, so treat every figure as a starting point checked in August 2026, not a quote. Second, the smart comparison is cost per job, not sticker price: a $29 keyword tool you use daily is better value than a $139 suite you open twice a month. Verify the current number on the vendor's own page before you commit.
ToolBest jobStarting price (Aug 2026)Note
SemrushBroadest all-in-one$139/mo ($117.33 annual)Closest all-around rival; AI Visibility a $99/mo per-domain add-on or Semrush One from $199/mo; 7-day trial
SE RankingCheaper all-in-one$129/mo ($103 annual)Basic AI-search visibility in-plan, fuller toolkit ~$79/mo add-on; white-label reporting is the Agency Pack add-on (+$69/mo); 14-day trial, no card
MangoolsKeyword research$29/mo (annual)AI Search Watcher included; free account + 48-hour money-back
UbersuggestCheapest entry$29/mo or $290 onceLifetime option; basic audits and AI Search Visibility bundled
MajesticBacklinks on a budget$49.99/moTrust Flow / Citation Flow; link-only
Moz ProDomain Authority + local$99/mo ($79 annual)DA metric and link explorer; 30-day trial, one of the longest
SurferContent optimization~$99/moOn-page scoring, not a full suite
SpyFuCompetitor PPC data$39/moAd history + estimated spend; US-heavy
Screaming FrogTechnical audits$279/yearFree for crawls under 500 URLs
Free / DIY stackOwn-site reporting$0Search Console + Analytics + Keyword Planner
geotoolbox (our tool)AI / GEO visibility$99/mo ($79 annual)Up to 8 engines by tier; also checks crawler access; 7-day trial
Ahrefs (staying put)Backlinks$29/mo StarterMonthly-only, no trial; Lite $129/mo for real research; Brand Radar AI from $199/mo per platform (Select), $699/mo All Platforms
The takeaway is simple: almost every single-job alternative costs a fraction of the suite. The only reason to pay suite prices is that you genuinely need most of the suite, or that backlinks are the job and you want the freshest, highest-quality index. ## How to Choose: Match the Tool to the Job You do not need a single winner. Match the tool to the job in front of you, and build a small stack only if you genuinely work across more than one. Moz Pro is worth a look too if the Domain Authority metric and a long 30-day trial matter to you. And if the tool you are actually comparing against is Semrush rather than Ahrefs, our [Semrush alternatives](https://geotoolbox.ai/blog/semrush-alternatives) guide runs the same job-based logic from that starting point.
If you are a...Start withWhy
Solo or small siteFree stack, then Mangools or UbersuggestOne paid tool for your main job; keyword research at a fifth of the suite price
Agency or in-house teamSemrush or SE RankingSuite breadth plus PPC and client reporting; SE Ranking is the cheaper all-in-one
Link builderStay on Ahrefs (or Majestic on a budget)Freshest, highest-quality backlink index; do not trade this down
Content-focusedSurfer or ClearscopeGrades drafts against what ranks; pair with a keyword tool
AI-visibility-focusedgeotoolbox (our tool) or a suite add-onTrack your mentions and citations in the answer engines, and check they can crawl you; do not buy a second full suite
Whatever you land on, if you start a trial, calendar the end date and cancel a few days early if you are not converting. That one habit defuses the billing surprises that send most people looking for alternatives in the first place. ## Frequently Asked Questions ### Is there a free alternative to Ahrefs? For your own site, the free stack of Google Search Console, Google Analytics, and Keyword Planner covers most of what people open Ahrefs for. For keyword ideas specifically, the Keyword Surfer extension and Ubersuggest's limited free tier help, and Ahrefs offers free Webmaster Tools for sites you verify you own. None replace paid competitive research, but they are enough to run a small site. ### What are the cons of Ahrefs? The three that drive people to alternatives are price (the Starter plan is limited, and real research starts at $129 a month), credit metering (report, crawl, and API credits that burn fast in a deep session), and the lack of a proper free trial. It is also thinner than Semrush on PPC, local, and content tooling, and its AI-visibility tracking is a costly add-on rather than a built-in feature. ### What is the cheapest Ahrefs alternative? Among paid tools, the lowest entry points are Mangools and Ubersuggest at $29 a month (Ubersuggest also has a $290 lifetime option), with Ahrefs' own Starter plan also at $29 but limited to spot-checks. These are cheap entry tiers, not full-suite replacements. Below them, the free DIY stack costs nothing if you only need to report on your own site. ### Does Ahrefs have a free trial? No. Ahrefs does not offer a classic time-boxed free trial. It has Ahrefs Free (the renamed Webmaster Tools) for sites you own, but you cannot run a competitor's domain through it. SE Ranking offers a 14-day no-card trial, and Semrush a 7-day trial (a card is required, so calendar the end date). ### Can I use Ahrefs for free? Partly. Ahrefs Free gives you a handful of tools, but the ones most people want, Site Explorer and Site Audit, only work after you verify ownership of the site, so you cannot use it for competitor research. For that, you need a paid plan or a free-tier competitor. ### Is there an Ahrefs alternative that tracks AI or ChatGPT visibility? Yes. SE Ranking includes basic AI-search tracking in its plans (with a fuller toolkit as an add-on), Mangools and Ubersuggest bundle it into their plans, Semrush sells its toolkit as a $99-a-month per-domain add-on, and Ahrefs' Brand Radar covers it as a separate product. Dedicated tools like geotoolbox (our tool) are built specifically to track mentions and citations across ChatGPT, Perplexity, Google AI Overviews, and other engines, and also check whether those engines can crawl your pages, which most suite add-ons still do not. ### What is Ahrefs Brand Radar, and is there a cheaper alternative? Brand Radar is Ahrefs' AI-visibility product. It tracks how often your brand appears and gets cited in AI answers, sold as two standalone plans: Select Platforms from $199 a month per platform or All Platforms at $699 a month. That is a real line item, which is why a dedicated tool or a suite that bundles AI tracking (SE Ranking, Mangools) is usually cheaper if AI visibility is the job. ## Start With the Job, Not the Brand There is no single Ahrefs replacement. What there is, for almost every reason people leave, is a cheaper or more focused tool for the specific job: Semrush for the broadest suite, SE Ranking for the same shape at a lower price, Mangools or Ubersuggest for keywords, Majestic for budget backlinks, Surfer for content, Screaming Frog for audits, the free stack for your own reporting, and a dedicated tracker for AI visibility. If backlinks are the whole job, staying on Ahrefs is the honest answer. The choice gets easy once you name the job instead of the brand. The one job the SEO suites still treat as an add-on is whether AI answers can see and cite you, and it is one of the fastest-growing reasons to reassess your stack in 2026. If that is the part you want to check, our free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) takes about a minute and tells you whether the answer engines can even reach your pages, no subscription required. Start there, then decide what the rest of your stack actually needs to do. ## Sources - Ahrefs Pricing (Starter, Lite, Standard, Advanced; credits) - `ahrefs.com/pricing` - Ahrefs Brand Radar (Select Platforms and All Platforms pricing) - `ahrefs.com/brand-radar` - Semrush Pricing (SEO plan, Semrush One, AI Visibility Toolkit) - `semrush.com/pricing` - SE Ranking Pricing (Core, Growth; AI Search Toolkit) - `seranking.com/pricing` - Mangools Plans and Pricing (AI Search Watcher) - `mangools.com/plans-and-pricing` - Ahref's cheaper alternative? - r/SEO community thread - `reddit.com/r/SEO/comments/11xkp05/ahrefs_cheaper_alternative` --- ## Semrush Alternatives (2026): The Honest, Job-Based List > The best Semrush alternatives for 2026, picked by job: cheaper all-in-one, keyword research, backlinks, PPC, local, and AI-visibility tracking. - Canonical: https://geotoolbox.ai/blog/semrush-alternatives - Published: 2026-08-18 · Updated: 2026-08-22 If you are looking for Semrush alternatives, you have probably hit one of three walls: the price, the add-on tax, or the billing. The answer most listicles dodge is that there is no single drop-in replacement. Semrush bundles eight different jobs into one expensive login, and the right alternative depends on which of those jobs you actually do. So this guide is organized by job, not by a ranking that happens to put the highest-paying affiliate first. Every price here was checked in August 2026, we flag where we have a commercial interest, and we tell you where a free tool beats a paid one. Find the row that matches your problem and skip the rest. ## Why People Look for Semrush Alternatives Semrush is the most complete SEO platform most people can buy, and almost nobody leaves it because the data is bad. They leave for one of a few specific reasons, and knowing which one is yours tells you which alternative to pick. **The price, and the jump between tiers.** The entry SEO plan is $139 a month, and the number on the pricing page is rarely the number on your invoice. Single-seat plans mean extra users cost from about $45 a month each, and the tools that justify the price sit behind higher tiers. For a solo consultant or a small site, you are paying for dozens of tools to use five. **The add-on tax.** This is the 2026 version of the price complaint. Semrush's [AI Visibility Toolkit](/go/semrush-ai?ref=semrush-alternatives-why) is a separate purchase at $99 a month per domain on top of your SEO plan, and the community's blunt summary is that "everything is sectioned off with separate pricing for different features." A plan that reads as $139 can land near $280 once the AI layer and a second seat are on it. **The billing, which is the complaint most lists skip.** This is worth being specific about because it is the loudest theme in the forums, louder than any feature gap. Users report a two-step cancellation: you cancel, then confirm through a link in a follow-up email, and if that email lands in spam you can be charged anyway. Per Semrush's [refund policy](https://www.semrush.com/company/legal/refund-policy/), the money-back guarantee applies only to a first annual subscription, not to monthly plans. Enough people have been caught by a renewal that a recurring piece of forum advice is to dispute the charge with their bank rather than wait on support, which tells you how the experience lands even if it is not the path we would start with. Our own [Semrush review](https://geotoolbox.ai/blog/semrush-review) covers the exact cancellation mechanics if you want them before you start a trial. There is also a simpler reason: you only report on your own site. If you are checking your own rankings and traffic rather than running competitive research at volume, Semrush is the wrong shape of tool. Search Console and a free keyword tool cover most of what you actually open it for. And there is the ownership question. Adobe [completed its acquisition of Semrush](https://news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition) on April 28, 2026. Nothing has broken, but "what happens to pricing under Adobe" is a fair reason to know your options before a renewal. None of these means Semrush is bad. They mean the tool is a poor fit for a specific job, and there is almost always a cheaper, more focused option for that job. ## Why There's No One-to-One Semrush Alternative Semrush is not one product. It is keyword research, backlink analysis, site audits, rank tracking, PPC intelligence, local SEO, content optimization, and now AI-search visibility, bundled into one login. No competitor matches all of that, which is why "what's the best Semrush alternative?" has no single answer. The best alternative depends on which of those jobs you actually do. That also means switching has a cost, and the better lists say so. People who cancel and try to rebuild their stack from three cheaper tools often come back, and the reasons are consistent: the keyword database is thinner, the backlink index is smaller and slower to update, there is no local pack, and each tool has its own learning curve. One tool that does eight jobs adequately can beat three tools that each do one job well, once you count the hours of stitching them together. So the useful question is not "what replaces Semrush?" It is "which one thing am I actually paying Semrush for, and what does that one thing cost somewhere else?" The rest of this guide is organized around that question, one job at a time. ## Semrush Alternatives at a Glance

Almost every single-job alternative undercuts the suite. The last column is the one the SEO suites treat as an afterthought.

The reasoning behind each pick is below, and a fuller pricing table with the caveats is near the end. ## The Best Semrush Alternatives, by the Job You're Hiring For ### If You Want the Same Kind of Toolkit, for Less The closest thing to "Semrush but cheaper" is [SE Ranking](/go/seranking?ref=semrush-alternatives-seranking). It covers the same core jobs, keyword research, rank tracking, site audits, competitor analysis, and backlink monitoring, at a lower price than Semrush. Note one thing most listicles get wrong: its plans include basic AI-search tracking, but the fuller AI Search toolkit (across AI Overviews, AI Mode, Perplexity, and ChatGPT) is a paid add-on (about $79/mo) on top of the Core plan, so on serious AI tracking it is not meaningfully cheaper than Semrush's add-on. The reason to switch is the lower price on the core SEO suite, not free AI tracking. The Core plan runs $129 a month, and on annual billing it comes in under Semrush's $117.33 annual SEO plan on the core SEO suite. Report generation is in-plan, though white-label client reporting is a paid add-on (the Agency Pack, from $69/mo annual), and the 14-day trial does not ask for a card. The real saving over Semrush is modest on the base price, around $10 to $14 a month, and the AI-search layer costs extra on both, so switch for the cheaper suite and the no-card trial, not for free AI tracking. **Watch for:** the backlink index is smaller and less fresh than Ahrefs or Semrush, so if links are your main job, this is not the pick. SE Ranking's pricing also scales with how often you want rankings checked, so a high-frequency setup costs more than the sticker. **Our pick for most people leaving Semrush:** [SE Ranking](/go/seranking?ref=semrush-alternatives-verdict) covers the same core jobs and undercuts Semrush on annual billing (its AI-search tracking is a separate add-on, like Semrush's). The 14-day trial does not ask for a card, so you can run your own site through it before deciding. ### If You Just Need Keyword Research For keyword work without the suite, [Mangools](/go/mangools?ref=semrush-alternatives-mangools) is the long-standing favorite. Its KWFinder tool is genuinely pleasant to use, the difficulty scores are reasonable, and the Basic plan is $29 a month billed annually, roughly a fifth of Semrush's entry price, and it is now sold as an "AI + SEO" bundle that includes AI Search Watcher. There is no free trial anymore, but a free limited account and a 48-hour money-back guarantee. **Watch for:** it is a keyword and SERP tool, not a full platform. The backlink and audit tooling is shallow, so treat it as a focused instrument rather than a Semrush replacement. Neil Patel's Ubersuggest is the other cheap option here, at $29 a month or a one-time $290, and it is the better fit if you want basic audits bundled with keyword ideas. ### If Backlinks Are the Whole Job When the job is links, most people do not trade down from Semrush, they trade up to Ahrefs. Its backlink index is widely regarded as the deepest and most frequently updated of the major tools, and its data is what the rest of the field gets compared against. If you spend your day on link research, disavows, and competitor backlink gaps, this is the tool. Ahrefs offers a $29-a-month Starter plan, a real entry point for spot-checks, though it is billed monthly only with no free trial. Serious research still means the Lite plan at around $129 a month. We put the two suites head to head on our own sites in [Semrush vs Ahrefs](https://geotoolbox.ai/blog/semrush-vs-ahrefs) if you are choosing between them specifically, and if you are looking to move off Ahrefs rather than toward it, our [Ahrefs alternatives](https://geotoolbox.ai/blog/ahrefs-alternatives) guide covers the options by job. For backlink data on a tighter budget, Majestic is the link-only specialist, built around its Trust Flow and Citation Flow metrics, from about $50 a month. **Watch for:** Ahrefs is weaker than Semrush on PPC data, social, and local, so it is a specialist's tool, not a broader marketing platform. ### If You Need Competitor PPC Data Semrush's paid-search intelligence is the reason a lot of agencies stay, but it is not the only place to get it. SpyFu is the budget specialist for competitor PPC and keyword research: it surfaces the keywords a rival buys and ranks for, their ad history, and estimated spend, from $39 a month ($29 billed annually). For a solo marketer who needs competitor ad data more than a full suite, it is a fraction of the cost. **Watch for:** SpyFu's data skews toward larger, US-heavy advertisers, and its organic and backlink coverage is thinner than the big suites. Use it for competitor ad intelligence, not as your only SEO platform. ### If You Only Report on Your Own Site Here is the option most tool roundups leave out: for watching your own rankings and traffic, you may not need a paid tool at all. Google Search Console shows the queries you actually rank for and the clicks you actually get, Google Analytics covers behavior, Keyword Planner gives volume ranges, and Looker Studio ties it into a report. That stack is free and, for your own site, more truthful than any third-party estimate. Paid suites earn their price on competitive research, the data about sites you do not own. If you are not doing that at volume, you are renting a race car to drive to the shop. Our [best free SEO tools](https://geotoolbox.ai/blog/best-free-seo-tools) guide maps the free stack job by job. Come back to a paid tool when competitor research, not self-reporting, becomes the work. ### If Content Optimization Is the Job If what you actually open Semrush for is its Writing Assistant, a dedicated content tool does it better. [Surfer](/go/surfer?ref=semrush-alternatives-surfer) scores a draft against the pages already ranking and gives concrete term and structure targets, at around $99 a month for its standard plan. Clearscope is the pricier, more editorial-grade option that in-house content teams tend to prefer. **Watch for:** these are on-page optimization tools, not research suites. They tell you how to shape a page, not which pages to build. Pair one with a keyword tool rather than expecting it to replace the whole platform. Our [content optimization tools](https://geotoolbox.ai/blog/best-content-optimization-tools) roundup compares the field. ### If You Live in Technical Audits or Rank Tracking For technical SEO, Screaming Frog is the standard desktop crawler and it is cheap for what it does: $279 a year for the full version, free for crawls under 500 URLs. It goes deeper on on-page and crawl issues than any suite's site audit, at the cost of a steeper learning curve. For rank tracking on its own, dedicated trackers like AccuRanker and Wincher are faster and cheaper than paying a full suite just to watch positions. If your tracking now includes AI answers as well as classic search, our [AI rank tracker](https://geotoolbox.ai/blog/ai-rank-tracker) guide covers what that shift changes. ### If Local SEO Is the Job None of the big suites are built for local first. If your work is Google Business Profiles, local pack rankings, and citations, a specialist like BrightLocal covers the job from $39 a month (about $31 billed annually), and Moz retains a reputation for local and its Domain Authority metric. Treat local as its own tool line rather than a reason to buy a whole platform. **Watch for:** BrightLocal's cheapest tier tracks citations and rankings but does not build citations for you, and reputation-management features sit on higher plans. ### If You Need Traffic and Market Intelligence One of the jobs people actually use Semrush for is estimating a competitor's traffic and sizing a market, and the tool built for that is [SimilarWeb](https://www.similarweb.com/). It estimates any site's traffic, sources, and audience, and benchmarks whole industries, which is why analysts and business-development teams reach for it more than SEOs do. Pricing is quote-based and climbs into enterprise territory quickly, so it is overkill unless market intelligence is the actual deliverable. **Watch for:** SimilarWeb is a traffic and market tool, not an SEO suite. It will not do keyword research, site audits, or rank tracking, and its self-serve pricing is opaque. ### The Job the SEO Suites Bury: AI and GEO Visibility Here is the disclosure up front, because it matters for this section: we build one of these tools. geotoolbox is our AI-visibility product, so read this as an interested party, not a neutral referee. The reason it belongs on a Semrush-alternatives list at all is that the AI-visibility job is the one the classic suites treat as an afterthought. Semrush answers it with a $99-a-month-per-domain add-on layered on top of your SEO plan; SE Ranking sells it as a paid add-on much like Semrush, while Nightwatch folds a version into its standard tiers; most of the field ignores it. If the thing you actually want to know is whether ChatGPT, Perplexity, Google's AI Overviews, and the rest mention and cite you, you are usually paying for a whole SEO suite to reach one module. A tool built for that job does two things a bolt-on rarely does well. It tracks mention and citation rates across the answer engines your buyers actually use, and it checks whether those engines can even fetch your pages, because a blocked AI crawler quietly sinks your visibility no matter how good the content is. [geotoolbox](https://geotoolbox.ai/pricing) starts at $99 a month ($79 annual, 7-day trial) and reaches up to eight engines on higher tiers, including Claude, Copilot, and Grok, which the self-serve suite add-ons did not cover as of mid-2026. That is also the real difference from SE Ranking's or Semrush's own suite modules: they track fewer engines and do not check whether those engines can actually crawl your pages. You can check the reachability half for free before paying for anything: our [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) tells you in about a minute whether those engines can actually fetch your pages, no account needed. Our [AI visibility tools roundup](https://geotoolbox.ai/blog/best-ai-visibility-tools) compares the whole field, including us, with the same disclosure made here.
![geotoolbox free AI-readiness scan result for developers.cloudflare.com showing a 100% grade-A score and passing core checks: robots.txt valid format, AI Bot Discoverability with 0 of 34 AI crawlers blocked, and sitemap present and valid.](/blog/semrush-alternatives/geotoolbox-ai-readiness.png)
The reachability half, free: our AI-readiness scan checks whether the answer engines' crawlers are actually allowed to fetch your pages before you pay to track visibility.
**Watch for:** geotoolbox does one job. It is not a Semrush alternative for the other seven jobs on this page, it has no keyword database, backlink index, or PPC data of its own, and it is a newer, smaller tool than the suites. And treat every AI-visibility number from any vendor, ours included, with some skepticism. These tools sample nondeterministic engines, so a "share of voice" figure is an estimate with a margin, not a meter reading. Buy on whether a tool measures the engines you care about and checks crawler access, not on a single headline percentage. If you are unsure the category is worth a line item yet, read [what AI visibility actually is](https://geotoolbox.ai/blog/what-is-ai-visibility) first. ### When Semrush Is Still the Right Call Not every reason to shop around survives contact with the alternatives. If you run an agency or a marketing team that will genuinely use three or more of its toolkits, Semrush is often still the right buy, and this guide would be dishonest to pretend otherwise. Its PPC competitive intelligence is deeper than anything else in a full suite (SpyFu is the budget alternative if that is your only reason to stay), the white-label client reporting is what the platform is built around, and having SEO, paid, content, local, and AI visibility in one login beats stitching five vendors together once you account for the coordination. If that is you, the smart move is to run the [7-day free trial](/go/semrush-seo?ref=semrush-alternatives-recapture) on your busiest account and decide on your own data, then calendar the trial end date so the billing does not decide for you. Our full [Semrush review](https://geotoolbox.ai/blog/semrush-review) walks through who should buy it and who should not. ## A Pricing Reality Check Two cautions before the table. First, these prices move constantly and often differ by region and billing period, so treat every figure as a starting point checked in August 2026, not a quote. Second, the smart comparison is cost per job, not sticker price: a $30 keyword tool you use daily is better value than a $139 suite you open twice a month. Verify the current number on the vendor's own page before you commit.
ToolBest jobStarting price (Aug 2026)Note
SE RankingCheaper all-in-one$129/mo (less on annual billing)Report generation in-plan, white-label client reporting an Agency Pack add-on (+$69/mo); basic AI-search in-plan, fuller toolkit ~$79/mo add-on; 14-day trial, no card
MangoolsKeyword research$29/mo (annual)AI + SEO bundle (AI Search Watcher incl.); free account + 48-hour money-back
AhrefsBacklinks$29/mo StarterMonthly-only, no trial; Lite ~$129/mo for real research
MajesticBacklinks on a budget~$50/moTrust Flow / Citation Flow; link-only
SpyFuCompetitor PPC data$39/mo ($29 annual)Ad history + estimated spend; US-heavy
UbersuggestCheapest entry$29/mo or $290 onceLifetime option; 7-day trial
SerpstatBudget all-in-one$50/mo (annual)7-day trial; monthly costs more
Moz ProLocal + Domain Authority$99/mo ($79 annual)30-day free trial, one of the longest
SimilarWebTraffic + market intelligenceQuote-basedEnterprise-leaning; not an SEO suite
BrightLocalLocal SEO$39/mo (~$31 annual)Tracks citations; building + reputation cost more
SurferContent optimization~$99/moOn-page scoring, not a full suite
Screaming FrogTechnical audits$279/yearFree for crawls under 500 URLs
Free / DIY stackOwn-site reporting$0Search Console + Analytics + Keyword Planner
geotoolbox (our tool)AI / GEO visibility$99/mo ($79 annual)Up to 8 engines by tier; also checks crawler access; 7-day trial
SemrushStaying put$139/mo ($117.33 annual)AI Visibility a $99/mo per-domain add-on; One bundles from $199/mo; 7-day trial
The takeaway is simple: almost every single-job alternative costs a fraction of the suite. The only reason to pay suite prices is that you genuinely need most of the suite. ## How to Choose: Match the Tool to the Job You do not need a single winner. Match the tool to the job in front of you, and build a small stack only if you genuinely work across more than one.
If you are a...Start withWhy
Solo or small siteFree stack, then Mangools or SE RankingOne paid tool for your main job; go broad or narrow to fit budget
Agency or in-house teamSemrush or SE Ranking + specialistsSuite breadth earns its price once you use three-plus toolkits
Link builderAhrefs (or Majestic on a budget)Deepest, freshest backlink index
PPC-focusedSpyFuCompetitor ad data for a fraction of a suite
AI-visibility-focusedgeotoolbox or an add-onTrack and get cited by the answer engines; do not buy a second full suite
**Solo operator or small site (roughly $30 to $130/mo).** Start with the free stack for your own reporting, then add one paid tool for the job you do most. If that job is broad, SE Ranking gets you closest to a full suite for the price, cheaper than Semrush on annual billing. If it is narrow, a $30 keyword tool like Mangools is plenty. **Agency or in-house team.** This is where the suite earns its price. If you will use three or more of its toolkits, need PPC data, or want client-ready reporting, keep Semrush or run SE Ranking as the cheaper all-in-one. Add specialists, Ahrefs for links, Screaming Frog for audits, only where the suite's depth runs out. **Anyone weighing AI visibility in 2026.** Do not buy a second full suite for it. Add a focused AI-visibility layer, or turn on the add-on in the suite you already run, and judge it on whether it tracks the engines your buyers use. If the real question is whether to run any of this in-house, weigh [agency versus software](https://geotoolbox.ai/blog/geo-services-vs-software) before you buy. Whatever you land on, if you start a trial, calendar the end date and cancel a few days early if you are not converting. That one habit defuses the billing complaint that sends most people looking for alternatives in the first place. ## Frequently Asked Questions ### What is the best free Semrush alternative? For your own site, the free stack of Google Search Console, Google Analytics, and Keyword Planner covers most of what people open Semrush for. For keyword research specifically, Ubersuggest and Mangools have limited free tiers, and Ahrefs offers free Webmaster Tools for your own verified sites. None replace paid competitive research, but they are enough to run a small site. ### What is the cheapest Semrush alternative? Among paid tools, the lowest entry points are Ahrefs' $29-a-month Starter plan and Ubersuggest at $29 a month (or $290 once), with Mangools close behind at $29 a month billed annually. These are cheap entry tiers, not full-suite replacements. Below them, the free DIY stack costs nothing if you only need to report on your own site. ### Is there a Semrush alternative that tracks AI or ChatGPT visibility? Yes. SE Ranking includes basic AI-search tracking in its plans and sells the fuller AI Search toolkit as a paid add-on, Nightwatch adds LLM monitoring, and dedicated tools like geotoolbox are built specifically to track mentions and citations across ChatGPT, Perplexity, Google AI Overviews, and other engines. Semrush covers it too, but as a separate $99-a-month add-on rather than in the base plan. ### Is Semrush a Russian company? No. Semrush Holdings is an American company headquartered in Boston, and since April 2026 it has been a subsidiary of Adobe. It was founded in 2008 by Oleg Shchegolev and Dmitri Melnikov, who are of Russian origin, which is where the question comes from, but the company is US-incorporated and was publicly traded on the NYSE before the Adobe acquisition. ### Why did Adobe buy Semrush? Adobe completed its roughly $1.9 billion acquisition of Semrush in April 2026 to fold search and AI-visibility data into its customer-experience and marketing stack. For existing users, the practical question is whether pricing changes under Adobe, which is a fair reason to know your alternatives before a renewal. ### Is Ahrefs better than Semrush? It depends on the job. Ahrefs is widely regarded as having the deeper, fresher backlink index and a cleaner interface, so specialists focused on links and organic research often prefer it. Semrush is broader, with PPC, local, social, and AI-visibility tooling Ahrefs does not match. We tested both on our own sites in [Semrush vs Ahrefs](https://geotoolbox.ai/blog/semrush-vs-ahrefs). ### Can free tools really replace Semrush? For reporting on your own site, largely yes. For competitive research at scale, no. The free stack cannot show you a competitor's full keyword or backlink profile the way a paid index can, and that gap is exactly what you are paying a suite for. Use free tools until competitor research becomes the actual work. ### Who competes with Semrush? Semrush's main competitors are Ahrefs (the closest all-around rival, strongest on backlinks), SE Ranking and Serpstat (cheaper full suites), Moz (established, local-friendly), Mangools and Ubersuggest (budget keyword tools), SpyFu (competitor PPC), Screaming Frog (technical audits), and SimilarWeb (traffic and market intelligence). For the AI-search visibility layer specifically, the competition is a newer set of tools, including geotoolbox. There is no single rival that matches Semrush across every job, which is why the right competitor depends on the job. ### What is the most accurate SEO tool? No third-party tool is "accurate" in an absolute sense, because they all estimate. For your own site, the only ground truth is Google Search Console, which reports the clicks and positions you actually get rather than a modeled estimate. When we ran Semrush and Ahrefs against Search Console on our own sites, both over-reported traffic to different degrees, which we documented in our [Semrush vs Ahrefs](https://geotoolbox.ai/blog/semrush-vs-ahrefs) test. The practical rule: use a paid tool for relative comparisons against competitors, and Search Console for the truth about yourself. ### Does Semrush have a free trial? Yes. Semrush offers a 7-day free trial that requires a credit card, and a limited free plan that covers about ten reports a day with no card. Because cancellation runs through an emailed confirmation link, calendar the trial end date if you start one. You can [start the trial](/go/semrush-seo?ref=semrush-alternatives-trial) to run your own data through it before deciding. ## Start With the Job, Not the Brand There is no single Semrush replacement. What there is, for almost every reason people leave, is a cheaper and more focused tool for the specific job: SE Ranking for the whole suite at a lower price, Mangools or Ubersuggest for keywords, Ahrefs for links, Screaming Frog for audits, the free stack for your own reporting, and a dedicated tracker for AI visibility. The choice gets easy once you name the job instead of the brand. The one job the SEO suites still treat as an add-on is whether AI answers can see and cite you, and it is the fastest-growing reason to reassess your stack in 2026. If that is the part you want to check, our free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) takes about a minute and tells you whether the answer engines can even reach your pages, no subscription required. Start there, then decide what the rest of your stack actually needs to do. ## Sources - Adobe Completes Semrush Acquisition - Adobe Newsroom, April 2026 - `news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition` - Semrush - Wikipedia (company, founders, headquarters) - `en.wikipedia.org/wiki/Semrush` - Semrush Pricing (SEO plans, One bundles, AI Visibility Toolkit) - `semrush.com/pricing` - Ahrefs Pricing (Starter, Lite, Standard) - `ahrefs.com/pricing` - SE Ranking Pricing (Core, Growth) - `seranking.com/pricing` - Mangools Plans and Pricing - `mangools.com/plans-and-pricing` - Moz Pro Pricing - `moz.com/products/pro/pricing` - Serpstat Pricing - `serpstat.com/pricing` - Surfer Pricing - `surferseo.com/pricing` - Screaming Frog SEO Spider (license + free version) - `screamingfrog.co.uk/seo-spider` --- ## AI SEO: What It Actually Means (and How to Do It in 2026) > AI SEO means two different things: using AI to do SEO faster, and optimizing to get cited by AI search. Here's the breakdown, with 2026 data. - Canonical: https://geotoolbox.ai/blog/ai-seo - Published: 2026-08-16 · Updated: 2026-08-17 "AI SEO" gets used for two completely different jobs, and most articles never say which one they mean. One is using AI tools to do ordinary SEO faster. The other is optimizing so that AI search engines like ChatGPT, Perplexity, and Google's AI Overviews mention and cite you. They need different work and different metrics. This guide separates the two, then answers the question underneath most of the confusion: is any of this new, or is "AI SEO" just regular SEO with a markup? The short version is that the craft barely changed, but the goalposts did.
![The two meanings of AI SEO side by side: Meaning 1, using AI to do SEO (AI-assisted SEO) - goal is doing existing SEO faster, metric is rankings and traffic; Meaning 2, optimizing for AI search (GEO/AEO/LLM SEO) - goal is getting mentioned and cited in AI answers, metric is mention rate, citation rate, and share of voice.](/blog/ai-seo/two-meanings-of-ai-seo.png)
"AI SEO" covers two different jobs with two different scorecards, and the confusion comes from treating them as one.
## What Is AI SEO? AI SEO is an umbrella term for two related but distinct practices: 1. **Using AI to do SEO** (also called AI-assisted SEO, or "SEO with AI"): pointing large language models and machine-learning tools at your existing SEO work so it goes faster. Keyword clustering, content briefs, technical audits, internal-link suggestions, content refreshes. 2. **Optimizing to be found by AI search** (usually called [generative engine optimization (GEO)](https://geotoolbox.ai/blog/what-is-geo), [answer engine optimization (AEO)](https://geotoolbox.ai/blog/what-is-answer-engine-optimization), or [LLM SEO](https://geotoolbox.ai/blog/llm-seo)): shaping your content and your wider web presence so AI answer engines discover, trust, and cite you as a source. The first is a workflow question. The second is a visibility question. Here is the split at a glance:
 Using AI to do SEOOptimizing for AI search (GEO/AEO)
GoalDo existing SEO work faster and cheaperGet mentioned and cited inside AI answers
Typical workKeyword clustering, drafts, audits, refreshes, schemaReachability, extractable structure, entities, being the primary source
Main riskHallucinated facts, mass-produced thin contentInvisible to the engines, or cited by them without any traffic
Primary metricRankings, organic traffic, conversionsMention rate, citation rate, share of voice
Which one you need depends on your problem. If your team is drowning in production work, you want meaning one. If you have noticed ChatGPT recommending competitors and never you, you want meaning two. Most serious operators end up doing both. ### What "SEO for AI" Is Called Now If you have seen a dozen acronyms and assumed they are competing standards, they are mostly the same idea seen from different angles. [GEO vs AEO vs SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) breaks the distinctions down, but in practice: **AEO** leans toward giving direct answers to specific questions, **GEO** and **LLM SEO** lean toward getting synthesized and cited by generative models, and "AI search optimization" is the plain-language catch-all. Google's own AI Overview for "ai seo" (observed August 2026) defined the term as both automating SEO tasks and optimizing for generative engines, a sign that the single label now covers both jobs. ## Is AI SEO Just Rebranded SEO? Mostly yes on the craft, genuinely no on the goal. This is worth being clear about, because a lot of what gets sold as "AI SEO" is a fundamentals package wearing a new label, and the fastest way to spot a weak pitch is that it cannot tell you what is actually different. Here is what did not change. The work that makes you visible to AI engines is, to a first approximation, the work that already made you a good search result: helpful content, clear structure, [semantic depth](https://geotoolbox.ai/blog/semantic-seo), real [entity signals](https://geotoolbox.ai/blog/entity-seo), and the [experience and expertise](https://geotoolbox.ai/blog/eeat-ai-search) that make a page trustworthy. If your SEO was sloppy, no amount of "GEO" fixes it. Rand Fishkin put the tension well when he argued that "your SEO still matters as much or more than ever before, it just won't earn you traffic the way it once did." Three things did change, and they are what make the label more than marketing: - **The objective.** The old goal was to rank so a person clicks. The new goal is to be the answer, or the source behind it, whether or not anyone clicks through. - **The unit.** Traditional SEO optimizes the page. AI engines pull passages and claims. The winning unit is now a self-contained, quotable statement, not a whole document. - **The measurement.** Rankings and clicks miss most of what is happening inside an AI answer, so the scorecard has to change too. So when a vendor pitches "AI SEO" as a bolt-on retainer, the fair question is which of those three they are actually addressing. If the answer is "we will use ChatGPT to write more blog posts," that is meaning one dressed up as meaning two, and on its own it does not amount to a visibility strategy. ## What Actually Changed: the Search Data The reason AI SEO exists as a category is that an answer layer now sits between your page and the searcher, and it is absorbing the click you used to earn. The numbers are stark. [Pew Research](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) analyzed 68,879 searches from 900 U.S. adults in March 2025. About one in five searches produced an AI summary, and 88% of those summaries cited three or more sources. But when an AI summary appeared, users clicked a traditional result only 8% of the time, versus 15% without one. They also left more often: 26% ended their browsing session right after a page with an AI summary, compared with 16% on a normal results page. Zoom out and the trend is the same. [SparkToro's clickstream analysis](https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/) found that 68.01% of U.S. Google searches ended without a click to an outside site in early 2026, up from 60.45% in 2024. The pure AI platforms barely link out at all: SparkToro estimates they send less than 1% of their traffic back to the web. Two conclusions follow, and they are the whole reason the discipline splits from classic SEO: 1. **Being seen no longer means being visited.** An AI Overview can put your name in front of someone who never clicks. That exposure can still be worth something (brand, trust, influence on the answer itself), but your analytics will not show it as traffic. 2. **You have to optimize for the answer, not just the ranking.** This is the same shift we cover in [is SEO dead in 2026](https://geotoolbox.ai/blog/is-seo-dead) and, mechanically, in [how AI search actually works](https://geotoolbox.ai/blog/how-does-ai-search-work): one query gets [fanned out](https://geotoolbox.ai/blog/query-fan-out) into many, and the engine assembles an answer from whatever sources it trusts. None of this means SEO stopped working. It means the payoff is shifting from the click toward the citation, and your strategy has to follow it. ## Using AI to Do SEO: What Works, What's Risky This is the meaning most people reach for first, and it is genuinely useful. AI is good at the high-volume, low-judgment parts of the job: clustering a messy keyword list, drafting content briefs, summarizing what the top-ranking pages cover, generating FAQ answers and schema, and flagging pages that have gone stale. That frees you to spend time on the parts that actually differentiate a page. Where it goes wrong is when people hand it the judgment too. Two failure modes show up again and again: - **Hallucination.** Models state wrong facts with total confidence, invent statistics, and cite studies that do not exist. On anything a reader will act on, every AI-supplied fact needs a human check against a real source. - **Scaled thin content.** It is trivial to generate a thousand near-identical pages. It is also a reliable way to get demoted. That second risk raises the question people ask most: **does Google penalize AI-generated content?** The plain answer is no, not for being AI-made. Google's [own guidance](https://developers.google.com/search/blog/2023/02/google-search-and-ai-content) says it rewards high-quality content "however it is produced," and that appropriate use of AI is not against its rules. What it does target, under [its spam policies](https://developers.google.com/search/docs/essentials/spam-policies), is "scaled content abuse": using automation to produce many pages primarily to game rankings rather than help people. The operative phrase is "without adding value," not "using AI." The practical rule that keeps you on the right side of this is a clean division of labor. Let the machine handle volume, and keep a human on judgment:
Let AI do itKeep a human on it
Keyword clustering and groupingDeciding the angle and what to leave out
First-draft outlines and FAQ structureFirsthand experience, original data, real examples
Parsing competitor pages for gapsFact-checking every claim before it ships
Generating schema and meta tagsVoice, editorial standards, and the final read
If you want the full workflow for the visibility side rather than the production side, our [step-by-step playbook for optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) goes deeper than we can here. ## Getting Cited by AI Engines: the Real Work This is the harder, newer half, and it has a prerequisite almost every guide skips: the AI engines have to be able to fetch your pages before any of the clever optimization matters. When we audit sites for AI visibility, one of the most common problems we find is not weak content, it is that the [AI crawler](https://geotoolbox.ai/glossary/ai-crawler) that fetches pages for answers was blocked before it could read the page, usually by a robots.txt rule the owner never meant to apply to it. The bot that matters here is the retrieval crawler each engine uses to build answers, like [OAI-SearchBot](https://geotoolbox.ai/glossary/oai-searchbot) or PerplexityBot, not the separate training crawler such as [GPTBot](https://geotoolbox.ai/glossary/gptbot). Reachability comes first. (And no, an [llms.txt file](https://geotoolbox.ai/blog/llms-txt) is not the fix; there is still no evidence it earns you citations.) Once the engines can read you, the work is about being the thing they want to quote: - **Extractable structure.** Clear headings, and passages that answer one question completely on their own, so a model can lift a self-contained chunk without losing the meaning. - **Entities and authority.** Consistent, well-defined [entities](https://geotoolbox.ai/blog/entity-seo) and [schema markup](https://geotoolbox.ai/blog/schema-markup-for-ai) that tell engines who you are and what you are an authority on. - **Being the source, not a summary.** Original data, firsthand testing, and specific numbers are more likely to get cited. Content that only restates what is already on the web is easy for a model to pass over in favor of the page that said it first. The part that makes this a real discipline rather than a rebrand: the surfaces do not agree with each other. [Ahrefs found](https://ahrefs.com/blog/search-rankings-ai-citations/) that in July 2025, 76% of the pages cited in Google's AI Overviews also ranked in Google's traditional top 10, but [a March 2026 update](https://ahrefs.com/blog/ai-overview-citations-top-10/) across 863,000 SERPs cut that to 38%: even on the surface closest to classic SEO, ranking well counts for less than it used to. But ChatGPT and Perplexity draw from very different source sets: [one analysis of 680 million AI citations](https://www.5wpr.com/research/state-of-ai-citations-2026/) found only about 11% of domains are cited by both. In practice, ranking well on Google, getting cited by [ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt), and getting cited by [Perplexity](https://geotoolbox.ai/blog/perplexity-seo) are distinct outcomes you measure separately when those surfaces matter to your audience. The same is true for [getting cited in Gemini](https://geotoolbox.ai/blog/gemini-seo). ## How to Measure AI SEO If most citations never produce a click, your old dashboard is blind to the thing you are now optimizing for. Rankings and organic traffic still matter, but they miss the answer layer entirely. The fix is a second set of metrics built around presence in AI answers rather than position on a page.
MetricWhat it measuresHow to track it
Mention rateHow often you appear in AI answers for prompts you care aboutRun a fixed prompt set across engines on a schedule
Citation rateHow often you are the linked source, not just namedSame prompt set, count linked citations
Share of voiceYour mentions versus competitors for the same promptsTrack the same prompts for rival brands
AI referral sessionsActual visits from AI surfacesGA4 traffic from chatgpt.com, perplexity.ai, gemini.google.com
Branded-search liftPeople who saw you in an answer and searched you laterSearch Console brand queries and direct traffic trend
The mention, citation, and share-of-voice numbers are the ones that map onto everything above: they tell you whether you are being quoted at all, even when nobody clicks. The referral and branded-search lines are how you connect that visibility back to revenue, which matters because last-click attribution will always undercount an answer that never sent a click. We go deeper on the mechanics in [how to measure your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility), the [AI visibility score](https://geotoolbox.ai/blog/ai-visibility-score) that rolls these into one number, and [tracking brand mentions in AI search](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search). ## Do You Actually Need AI SEO? Not every business needs to act on this today. The deciding question is whether your audience has started using AI to research what you sell. You should be working on AI visibility now if your buyers ask research-heavy or comparison questions (most B2B, software, finance, health, and considered purchases), because those are exactly the queries AI engines answer directly. You can afford to wait, and lean on [classic SEO plus the budget split we lay out in GEO vs SEO](https://geotoolbox.ai/blog/geo-vs-seo), if you win mostly on local or transactional intent where people still click through to buy, book, or call. The bottom line: for most brands the right move in 2026 is not to pour a budget into "AI SEO" as a separate line item, it is to keep doing strong SEO and add the visibility work and the measurement above on top. If you would rather hand it off, compare your options with our guides to [AI SEO agencies](https://geotoolbox.ai/blog/best-ai-seo-agencies) and [AI visibility tools](https://geotoolbox.ai/blog/best-ai-visibility-tools) before signing anything, and hold any vendor to the "what is actually different" test from earlier. ## Where to Start If all of this feels like a lot, start with the cheapest, highest-impact move: find out whether AI engines can even see you. Before any content strategy, before any tooling budget, a page a retrieval crawler cannot fetch is very unlikely to be cited, no matter how good it is. That is the check we built geotoolbox to run. Our free [AI readiness tool](https://geotoolbox.ai/tools/ai-readiness) shows you which AI crawlers can reach your site and where they are being blocked, so you fix the invisible problem before you spend on the visible one. Get reachability right first, then work down the list above. ## Frequently Asked Questions ### What is SEO for AI called now? There is no single name yet. The common ones are generative engine optimization (GEO), answer engine optimization (AEO), LLM SEO, and AI search optimization. They overlap heavily and mostly describe the same goal: getting your content mentioned and cited by AI answer engines rather than just ranked in a list of links. ### Can ChatGPT do SEO? ChatGPT can do the tasks: keyword ideas, outlines, meta descriptions, drafts, and audits of existing pages. It cannot do the judgment reliably, and it invents facts, so anything it produces needs a human check. It also cannot reliably tell you what has real search demand, since it has no keyword-volume or rank-tracking data. Treat it as a fast assistant, not a strategist. If your goal is the reverse, getting cited inside ChatGPT, see our guide to [SEO for ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt). ### Does Google penalize AI-generated content? No, not for being AI-generated. Google's guidance says it judges content on quality and helpfulness however it was produced. What it penalizes is scaled content abuse: mass-producing low-value pages to manipulate rankings, whether a human or a model wrote them. Original, accurate, genuinely useful content is fine regardless of how you drafted it. ### Is AI SEO worth it for a small business? It depends on whether your customers use AI to research what you sell. If they ask AI assistants comparison or how-to questions in your category, being cited is worth pursuing now. If you win on local or transactional searches where people still click to buy, solid traditional SEO covers most of your upside and AI visibility can wait. ### Is SEO dead because of AI? No. The click you used to earn by ranking is shrinking fast, but the work of getting found is very much alive; it is just being paid out in citations and mentions instead of guaranteed clicks. We cover the data behind this in [is SEO dead in 2026](https://geotoolbox.ai/blog/is-seo-dead). ### Can I do AI SEO myself? Yes, especially the parts that matter most. Start with reachability, which is a free check anyone can run: confirm the retrieval crawlers can fetch your key pages. From there, the on-page work (clear structure, self-contained answers, real expertise) is the same craft as good SEO, and you can use AI tools to speed up the drafting. The judgment (what to say, whether a claim is true, your firsthand angle) is the part you keep for yourself, whether you hire help or not. ### How is AI SEO different from regular SEO? The craft is largely the same: helpful content, clean structure, and real [topical authority](https://geotoolbox.ai/blog/topical-authority). What changed is the objective (be the answer, not just rank), the unit (a quotable passage, not the whole page), and the measurement (mentions and citations, not just clicks and positions). ## Sources - Google users are less likely to click on links when an AI summary appears - Pew Research Center, July 22, 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` - In 2026, Less Than One Third of Google Searches Still Send a Click - SparkToro, 2026 - `sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click` - 76% of AI Overview Citations Pull From the Top 10 - Ahrefs, July 21, 2025 - `ahrefs.com/blog/search-rankings-ai-citations` - Update: only 38% of AI Overview citations come from the top 10 - Ahrefs, March 2026 (863,000-SERP study) - `ahrefs.com/blog/ai-overview-citations-top-10` - The State of AI Citations 2026 (680M-citation analysis) - 5WPR - `5wpr.com/research/state-of-ai-citations-2026` - Google Search's guidance about AI-generated content - Google Search Central, February 8, 2023 - `developers.google.com/search/blog/2023/02/google-search-and-ai-content` - Spam Policies for Google Web Search - Google Search Central - `developers.google.com/search/docs/essentials/spam-policies` --- ## Is SEO Dead in 2026? What the Data Actually Says > Is SEO dead in 2026? No - but 'rank #1, get the click' is. The data on AI Overviews, zero-click search, and what actually gets you cited. - Canonical: https://geotoolbox.ai/blog/is-seo-dead - Published: 2026-08-16 · Updated: 2026-08-16 Is SEO dead? No. But if you are waiting for the old deal, rank number one and collect the click, that part is genuinely dying, and the numbers back up the dread you are feeling. Search did not disappear. It split. The work of getting found is still very much alive, but the payoff is moving from the click you earn by ranking to the mention you earn by being citable. You can see the shift in the question itself: people are asking AI assistants whether SEO is dead far more than they were a year ago.
![Estimated monthly U.S. demand for the query 'is SEO dead' on AI assistants (DataForSEO index) rose from about 215 in January 2026 to 12,528 by July 2026.](/blog/is-seo-dead/ai-search-volume-surge.png)
The question itself is increasingly asked of AI assistants. Estimated monthly demand for "is SEO dead" on AI assistants rose from ~215 (Jan 2026) to ~12,500 (Jul 2026). Source: DataForSEO AI-search-demand estimates, pulled August 2026.
## The Short Answer: SEO Isn't Dead, the Click Is **SEO is not dead. What is dying is the reflex that ranking number one hands you a click.** Two things that used to move together, your position and your traffic, have come apart. A page can sit in the top three and still lose half its clicks to an AI summary parked above it. A page that ranks nowhere near the first screen can get quoted inside an AI answer. Position and payoff have decoupled, and once you see that, most of the "SEO is dead" panic resolves into a more useful question: how do you earn attention when the click is no longer guaranteed? The discipline is not ending; the job is changing. It is moving from "rank a page and harvest the click" to "be the source an answer engine reaches for." Generative engine optimization (GEO) is the name for that second job, and it sits on top of SEO, not in place of it. Everything below is the evidence. ## Why It Feels Dead: Rankings Hold, Traffic Falls The most common version of this complaint is not "I dropped in the rankings." It is "I am still number one and my traffic fell off a cliff anyway." That pattern is real, and if that is your situation, you are not imagining it or doing SEO wrong. Here is what changed. When Google shows an AI summary at the top of the results, people click far less. A [Pew Research Center study](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) of 68,879 real searches, published in July 2025, found users clicked a result just 8 percent of the time when an AI summary was present, versus 15 percent when it was not, and only 1 percent clicked a link inside the summary itself. The answer is delivered on the page; the trip to your site never happens. Zoom out and the same story holds at scale. [SparkToro's 2026 analysis](https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/) of US clickstream data found 68 percent of Google searches now end without a click at all, the [zero-click majority](https://geotoolbox.ai/blog/zero-click-searches). Not all of that is AI, and it is worth saying so plainly: maps, knowledge panels, ads, and on-SERP answers were eating clicks long before AI Overviews arrived. But AI summaries widened a gap that was already there. The click-through damage is measurable. [Seer Interactive's September 2025 study](https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-september-2025-update) found organic click-through rate on queries with an AI Overview fell roughly 65 percent, from 1.76 percent in June 2024 to 0.61 percent by September 2025, close to Seer's own "cut by nearly two-thirds." So the frustration is legitimate. The click leaked out of the funnel while your ranking stayed exactly where it was, and that is a different problem than "SEO stopped working." Some operators say this feels structural rather than like a normal algorithm cycle, and on the click economics, they are right. For how these summaries pick and place sources, see our breakdown of [Google AI Overviews](https://geotoolbox.ai/blog/what-are-google-ai-overviews). ## What Actually Died (and What Didn't) "SEO is dead" is too blunt. Something specific died, and confusing it for the whole discipline is what sends people into a panic. What died is commodity informational content: the page that existed only to answer a simple question a machine can now answer in one paragraph. If your traffic came from "what is X" or a thin definition post, an AI summary now resolves that query without a click, and it is not coming back. Publishers of that kind of content report the steepest declines, some watching once-huge informational sites shrink to a fraction of their old traffic. That is the part of SEO that actually died. What did not die is anything an answer box cannot replace. Bottom-of-funnel and transactional pages still convert, because the searcher needs to do something, not just know something. High-trust and local queries still send clicks, because a person hiring an accountant or a photographer has to vet the actual business, not read a summary about the concept. Original data, first-hand experience, and truly unreplicable content still earn attention, because a model cannot regurgitate what only you have.
Losing ground fastHolding or growing
Commodity "what is / how to" explainers a summary now answersBottom-of-funnel and transactional pages (the searcher needs to act)
Thin listicles and definition posts written for a snippetLocal and high-trust queries where people vet a real business
Keyword-density pages with no original angleOriginal data, research, and first-hand experience
Volume-play content that only restated common knowledgeBrand and entity authority that answers get built around
This is where the "is SEO still relevant" question actually lands. It is relevant exactly where it was always strongest: intent that ends in an action, trust that has to be earned, and information only you can provide. The commodity middle is what the machines took. ## The Real Shift: Ranking No Longer Equals Being Cited The mechanism underneath all of this is a split you can measure. **Ranking and getting cited have come apart.** [Ahrefs' tracking of AI Overview citations](https://ahrefs.com/blog/ai-overview-citations-top-10/) found the share that also rank in Google's organic top 10 fell from about 76 percent in its July 2025 study to 38 percent in its March 2026 report. Put the other way, roughly 60 percent of the sources an AI answer now quotes sit outside the top 10, and about a third rank nowhere in the top 100. That breaks the oldest reflex in SEO: earn authority, rank higher, collect the traffic. A Surfer analysis of roughly five million AI citations found domain authority barely correlates with being cited, close to zero across the AI surfaces it studied. Authority is not worthless. It still helps you rank, and on Google's own AI Overviews, which are built on the ranked index, ranking still feeds citation. It is on the pure-LLM surfaces, the ones that [assemble answers their own way](https://geotoolbox.ai/blog/how-does-ai-search-work), where the old signals matter least. There is a reason brand still counts even when domain authority does not. A page-level signal only pays off if that exact page is the one the model pulls. A brand or entity signal pays off whenever the model composes an answer about your category, because it is reaching for what it associates with the topic, not ranking a URL. The unit that gets cited is shifting from the page to the entity, which is what people mean when they say brand signals matter for AI. You can watch the split on the query that brought you here. As of mid-August 2026, the US Google results for "is seo dead" open with an AI Overview that cites a handful of sources, including a Reddit thread that is also the number-one organic result, while a Forbes article on the first page is not cited at all. Page-one presence alone did not guarantee a citation. Community and forum pages punch well above their ranking in these answers, which says something about what the systems reach for. AI results also vary by user and shift over time, so read this as a dated snapshot, not a fixed rule. And being cited is worth it. On informational queries, [Seer's 2026 data](https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update) puts a cited page at about 2.1 percent click-through against 0.9 percent when an Overview appears without citing you, its own "+120 percent" on the unrounded figures. Seer flags that the uncited figure is dragged down by one outlier account; strip it out and the gap narrows to roughly 2.1 versus 1.6 percent. Either way a cited page beats an uncited one, though both still trail the click-through those queries earn with no Overview at all. So the game is less "rank or die" than "get cited or get skipped," and that is a job you can work on. ## Has the Bottom Fallen Out? The Part Nobody Reports Almost every "SEO is dead" article ends at the cliff: click-through collapsed, the end. The data has a second chapter that rarely gets quoted. After the initial drop, the decline stopped. In [Seer's larger 2026 panel](https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update) of 53 brands, click-through on AI-Overview queries bottomed near 1.3 percent in December 2025 and rebounded to about 2.4 percent by February 2026, an 85 percent jump in two months (panel-wide, across all query types, so above the informational-only figure in the last section). That panel tracks a different, larger set than the September 2025 study above, so the absolute numbers are not directly comparable. Seer calls it a leveling out at a new normal; we would put it plainly: stabilization, not a return to pre-AI baselines, since 2.4 percent still sits below the 3.8 percent those same queries earn with no Overview at all. Two months is also too little to call a durable floor. The premise that Google itself collapsed does not hold either. [StatCounter](https://gs.statcounter.com/search-engine-market-share) still put Google at about 91 percent of search-engine share in mid-2026. Google did not lose its grip; its own results page just started answering more queries in place. And the much-hyped replacement, direct traffic from ChatGPT and other AI assistants, is still tiny for most sites: [Ahrefs measured it at 0.1 percent of referral traffic](https://ahrefs.com/blog/ai-traffic-research/) across 35,000 sites in early 2025, and 2026 estimates still put it well under one percent. It is growing fast, and much of AI's real effect shows up indirectly, as branded searches and direct visits after someone reads about you in an answer, which is exactly why the raw referral number understates it. Anyone telling you AI search already replaced Google traffic is selling something. So both things are true at once. The click economics did get worse and are not going back to 2020, while the channel stabilizes at a new level and a second, citation-based way of getting found opens up beside it. That is not a dead discipline. It is one mid-reinvention, which is a more annoying thing to manage, because the work changed rather than ended. ## So What Is Replacing SEO? The Acronyms, Defined Ask five articles what replaces SEO and you get five answers, which is why the question feels unanswerable. Part of the confusion is that the loudest "SEO is dead" case and the "SEO evolved" case often prescribe the same thing. When Inc. columnist Joe Procopio [declared SEO dead](https://www.inc.com/joe-procopio/seo-is-dead-according-to-google/91193051), his remedy was to build a direct audience and, echoing a Daily Mail editor he quoted, lean into branded search and unreplicable content. That is not the opposite of evolution. It is a description of it, minus the word. It is fair to be skeptical of the vocabulary here. A lot of what gets sold as GEO is repackaged SEO with a new invoice, and some of it is snake oil. So be precise about what is actually new: the fundamentals below, crawlable pages, clear structure, real authority, a well-defined entity, are the same SEO work they always were. The work is not new. The scoreboard is. You are now optimizing to be the retrieved and cited source, not only the ranked one, and if a "GEO service" cannot tell you which of those two it is changing, it is selling you the invoice. That is the frame for the new alphabet. None of these terms replace SEO; they are layers on that same foundation.
TermWhat it optimizes forWhere it playsRelationship to SEO
SEO (search engine optimization)Ranking a page so it earns the clickGoogle and Bing resultsThe foundation everything else stands on
GEO (generative engine optimization)Getting cited or mentioned inside AI-generated answersChatGPT, Perplexity, Google AI Mode and Overviews, GeminiA layer on top of SEO, not a swap
AEO (answer engine optimization)Being the direct answer to a questionFeatured snippets, voice, answer boxes, AI answersThe snippet-era ancestor of GEO
LLMO (large language model optimization)Shaping how models represent your brandThe underlying models themselvesOften used interchangeably with GEO
The short version: SEO gets you into the index an answer engine reads from; GEO is the work of being the thing it quotes. If you want the deeper taxonomy, our [GEO vs AEO vs SEO breakdown](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) is the triage page, and [GEO vs SEO](https://geotoolbox.ai/blog/geo-vs-seo) covers how to split a budget between the two. What none of them are is a reason to stop doing SEO. You cannot be cited from an answer engine that cannot crawl, read, and trust your page in the first place. ## What Still Works, and What's Just Folklore Most wasted GEO effort comes from confusing what actually holds up with what is folklore. The panic around "SEO is dead" has spawned a fresh crop of cargo-cult tactics, things that sound like GEO but do nothing. Start with what genuinely moves the needle. Models favor current, correct sources, so freshness and accuracy count for more than they used to. Answer engines lift self-contained passages, which puts a premium on clear headings, direct answers, and clean schema that make a page easy to extract. The strongest signal is often off your own page entirely: getting mentioned and cited across the web does more than anything you can publish about yourself. And the basics still apply, since a fast, crawlable, trustworthy site is one a model can actually use. That last part is [E-E-A-T for AI search](https://geotoolbox.ai/blog/eeat-ai-search), and it did not go anywhere. Now the folklore. The biggest one is [llms.txt](https://geotoolbox.ai/glossary/llms-txt), a proposed file that supposedly tells AI models how to read your site. It sounds official and takes an afternoon to feel productive about. [Ahrefs checked 137,210 domains](https://ahrefs.com/blog/llmstxt-study/) and found 97 percent of llms.txt files received zero requests in a month, and no major AI engine has committed to reading it as a retrieval or citation signal. We will say that against our own interest: geotoolbox ships llms.txt tools, and the data still says almost nobody consumes the file.
Holds up under the dataFolklore and busywork
Freshness, accuracy, and genuinely original informationChasing domain authority as if it buys citations (it barely correlates)
Extractable structure: clear headings, direct answers, clean schemaPublishing an llms.txt file and expecting AI traffic
Third-party mentions and digital PR (being cited elsewhere)Keyword density and word-count targets
Technical health: fast, crawlable, trustworthyMore commodity content to "feed the algorithm"
If you want the constructive version of the left column as a workflow, we wrote the [step-by-step playbook for optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) separately. The one-line summary: earn mentions, be extractable, stay accurate, and skip the rituals. ## Is SEO Still Worth It? The Measurement Problem The verdict, since the title asked a yes-or-no question: **yes, SEO is still worth investing in, as the foundation your AI visibility is built on, but only if you change what you measure.** Keep grading the channel on rankings and raw organic clicks and it will look like it is dying, because those are the two numbers AI broke. The value moved to citations, brand mentions, and the branded searches and direct visits that follow when someone reads about you in an answer. That is why it is so hard to defend to a boss or a board. The most painful version of "SEO is dead" is not technical, it is a marketing lead losing a budget fight because leadership read a headline. The reframe: blog clicks are down, but the brand is showing up in the answers people act on. The problem is that most teams cannot yet see that second half. [Semrush's 2026 AI Visibility Index](https://www.semrush.com/news/463141-semrush-releases-expanded-2026-ai-visibility-index-analyzing-126-million-ai-search-prompts/) reported that 45 percent of marketing leaders cannot accurately measure their brand's visibility in AI answers, and only 9 percent have tools to track all the relevant metrics across platforms. So change the scoreboard. When clicks go flat, these are the numbers that show whether SEO is still working.
Measure thisWhat it tells youWhere to read it
AI citation and mention rateWhether engines quote you, not just rank youAI-visibility tracking; manual prompt checks
Branded search and direct trafficDemand and "dark" visits created by answers you appear inSearch Console; GA4 direct and unassigned channels
Conversions, not clicksWhether the traffic you keep still paysGA4 or your CRM
Rank on high-intent, transactional queriesThe SEO that AI has not eatenYour rank tracker
Two honest caveats before you build that dashboard. AI answers are nondeterministic: they change from one run to the next, by user, location, and session, so any AI-visibility number, ours included, is a sampled estimate across many checks, not a rank you look up once. And a lot of AI's effect is unattributable by design, since people read an answer, never click, and arrive later as branded or direct traffic, so expect your own numbers to undercount it. Does the traffic that does arrive convert better? Vendors claim it does, sometimes dramatically, but no independent study has confirmed a durable multiplier, so treat those figures as marketing until someone replicates them. Be clear-eyed about the ceiling too: a doubled click-through on a cited page beats being skipped, but a citation is not a sale, and it has not been shown to pay back the way a ranked, clicked page once did. Pretending otherwise is just the old hype pointed in a new direction. One remedy is worth naming precisely because it does not route to anything we sell: the surest hedge is to not depend entirely on Google, or on being cited at all. An owned audience, an email list, a community, a channel of your own, is the one asset an answer engine cannot sit in front of. Beyond that, the work is measurement, and full disclosure, that is also our line of business: we build [tools that track AI brand mentions](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search), so treat our conclusion with the skepticism any vendor deserves. The third-party studies here are not ours, and the takeaway would hold if we sold nothing: start instrumenting whether AI engines can reach your pages, whether they cite you, and what branded search and direct traffic do afterward. Those are the numbers that tell you SEO is working when the old ones say it is dead. ## The Bottom Line SEO is not dead. The click you used to get for free is dying, and the discipline is splitting into two jobs: ranking to earn clicks, and being citable to earn mentions. Pretending nothing changed is how you lose. So is torching a channel that still holds up wherever intent, trust, and original information live. The next question is usually "so where do I even start?" Start by finding out whether the AI engines can see you at all, because a page a crawler cannot read is a page no answer can cite. Our free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) scans your site for the reachability and structure issues that keep you out of AI answers, and it is the fastest way to turn "is SEO dead" from a feeling into a to-do list. The obituary was premature; the job just changed. ## Frequently Asked Questions ### Is SEO still worth it in 2026? Yes, as the foundation your AI visibility sits on, but only if you change what you measure. Rankings and raw organic clicks now understate SEO's value because AI Overviews absorb many clicks. The payoff has shifted to citations, brand mentions, and the branded searches and direct visits that follow, so grade the channel on those instead. ### What is replacing SEO? Nothing is replacing it; a layer is being added on top. Generative engine optimization (GEO) is the work of getting cited inside AI answers, and it depends on the same crawlable, structured, trustworthy pages that SEO produces. Answer engines cannot quote a site they cannot read, so SEO remains the foundation the newer layer stands on. ### Is SEO dead now with AI? No. What died is commodity informational content that an AI summary can now answer without a click. Bottom-of-funnel pages, local and high-trust queries, and original content still earn attention. Google also still handles over 90 percent of search, so the audience did not leave; the results page just answers more questions itself. ### Will AI kill SEO? It is changing SEO, not killing it. Click-through on AI Overview queries fell sharply, then stabilized at a lower level, while a second path to visibility, being cited in AI answers, opened alongside it. The skill set shifts toward earning mentions and being extractable, but the underlying job of getting found is very much intact. ### Is SEO worth learning in 2026? Yes, if you learn the version that is growing rather than the one that is shrinking. Chasing rankings for commodity informational content is a dead end. The durable skills are technical health, content structure, entity and brand building, digital PR, and measuring visibility across both search and AI answers, and they transfer directly into generative engine optimization. That is why the field is changing rather than closing. ## Sources - Google users are less likely to click on links when an AI summary appears - Pew Research Center, July 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` - In 2026, less than one-third of Google searches still send a click - SparkToro, June 2026 - `sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click` - AI Overview impact on Google CTR, September 2025 update - Seer Interactive, November 2025 - `seerinteractive.com/insights/aio-impact-on-google-ctr-september-2025-update` - AI Overview impact on Google CTR, 2026 update - Seer Interactive, April 2026 - `seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update` - Google AI Overviews CTR shows early signs of recovery - Search Engine Land, April 2026 - `searchengineland.com/google-ai-overviews-ctr-recovery-study-475566` - Only 38% of AI Overview citations rank in the top 10 - Ahrefs, March 2026 - `ahrefs.com/blog/ai-overview-citations-top-10` - AI traffic research: how much traffic AI sends - Ahrefs, March 2025 - `ahrefs.com/blog/ai-traffic-research` - llms.txt study: 97% of files got zero requests - Ahrefs, June 2026 - `ahrefs.com/blog/llmstxt-study` - Domain authority and AI citations - Surfer, July 2026 - `surferseo.com/blog/domain-authority-impact-on-ai-citations` - Expanded 2026 AI Visibility Index (126M prompts) - Semrush, June 2026 - `semrush.com/news/463141-semrush-releases-expanded-2026-ai-visibility-index-analyzing-126-million-ai-search-prompts` - Search engine market share - StatCounter GlobalStats, 2026 - `gs.statcounter.com/search-engine-market-share` - AI search-demand estimates for "is seo dead" - DataForSEO AI Optimization API, pulled August 2026 - `dataforseo.com` - SEO Is Dead, According to Google - Inc. (Joe Procopio), May 2025 - `inc.com/joe-procopio/seo-is-dead-according-to-google/91193051` --- ## Topical Authority: How It Actually Works for Google and AI Search > Topical authority is how search and AI engines learn to treat your site as the expert on a subject. Here is how it really works, and whether AI search cares. - Canonical: https://geotoolbox.ai/blog/topical-authority - Published: 2026-08-16 · Updated: 2026-08-16 Topical authority is a signal that your site is the expert source on a subject, not just a page that happens to rank for one keyword. Build it and Google surfaces you for a wider range of queries, and AI engines are more willing to pull you into an answer. That is the promise every guide repeats. Most of them stop there. This goes further: what topical authority is versus the terms people confuse it with, how it feeds AI citations, whether it is even a confirmed ranking factor, and the precondition that makes the whole thing worthless if you skip it. ## What Topical Authority Actually Is **Topical authority is the degree to which search and AI engines treat your site as a go-to source on a specific subject, across the full range of related queries within it.** It is a relationship between your site and a topic, not authority over a single page or keyword. The important word is *earned*. Topical authority is an outcome, not a lever you pull. You do not "add" it the way you add a title tag. It accumulates when you cover a subject deeply, connect that coverage so engines can see it as one body of work, and back it with real expertise. That distinction matters because the term gets dismissed. Ask on r/SEO and you will find people calling topical authority "not a real thing" invented by influencers. They have half a point: there is no button labeled topical authority. But the underlying idea, that engines model how completely a site covers a subject, is well supported. The mistake is treating it as a tactic instead of the result of good tactics. The most cited practitioner framing comes from [Koray Tuğberk Gübür](https://www.holisticseo.digital/theoretical-seo/topical-authority/), who frames topical authority as **topical coverage multiplied by historical data**: how completely you cover the subject, times how long and consistently you have done it. Read as a rough model rather than a real equation, the multiplication makes a point: coverage with no track record, or a long history with shallow coverage, leaves you short either way. Coverage builds the map; time and consistency turn it into trust. For a working [definition of topical authority](https://geotoolbox.ai/glossary/topical-authority) and how it connects to [entity SEO](https://geotoolbox.ai/blog/entity-seo), the short version is this: coverage tells the engine what you are about, and consistency tells it you can be relied on to be right. ## Topical Authority vs Topic Clusters, Topical Maps, and Domain Authority These terms get used interchangeably, and that is the single biggest source of confusion. They are not the same thing: some are how you build authority, one is the result, and one is a different metric entirely.
TermWhat it isWhere it sits
Topical authorityThe outcome: engines recognize your site as an expert on a subjectThe result you are trying to earn
Topic clusterThe architecture: a pillar page plus supporting pages, interlinkedThe structure you build to earn it
Topical mapThe blueprint: the full set of subtopics and questions you plan before writingThe plan that comes first
Domain authority (DR / DA)A third-party estimate of whole-site backlink strengthA vendor metric, not Google's, and not topic-specific
Read it as a sequence. You draw a **topical map** to decide what to cover. You publish that map as a **topic cluster** so the pieces are connected. If the coverage is deep and credible enough, you earn **topical authority** as a result. **Domain authority** sits off to the side: it measures backlinks across your entire site and says nothing about whether you own a subject. A low-DR site regularly outranks a high-DR generalist inside a niche it has covered completely. The other neighbor is [semantic and entity SEO](https://geotoolbox.ai/blog/semantic-seo), which is the *how* underneath all of this: making the concepts on your pages legible to engines as entities in a [knowledge graph](https://geotoolbox.ai/glossary/knowledge-graph), so the relationships between your cluster pages are machine-readable rather than merely implied. You can have a tidy topic cluster and still earn no authority if the coverage is thin or the site has no track record. A cluster makes coverage easier for engines to follow, but structure on its own does not earn authority. ## How Topical Authority Works for AI Search Everyone says topical authority helps you get cited by ChatGPT and Perplexity. Fewer explain the mechanism, and the mechanism is what tells you where to spend your effort. AI engines answer through a retrieval pipeline. A prompt gets broken into several narrower sub-questions, a step called [query fan-out](https://geotoolbox.ai/blog/query-fan-out). The engine retrieves candidate passages from a search index for each sub-question, re-ranks them on relevance and quality, and writes an answer that cites the passages it used.
![A four-stage flow: one prompt fans out into sub-queries, retrieval pulls candidate passages from the index (where topical coverage acts), then re-ranking, then a cited answer.](/blog/topical-authority/where-topical-authority-acts.png)
Topical authority works mainly at the retrieval stage, where broad coverage decides how many of your pages enter the candidate pool an AI answer is built from.
The signals bundled under topical authority act at more than one of these stages, and it helps to keep them apart. At the **retrieval stage**, what matters is breadth: because one prompt fans out into many sub-questions, a site that covers a subject completely has a page that matches more of them, so more of your passages enter the candidate pool in the first place. At the **re-ranking and citation stage**, what matters is credibility: among the passages that were retrieved, engines lean toward sources they have reason to trust on the subject. Coverage gets you into the pool. Trust helps you get picked from it. That is why depth compounds. A single strong page answers one sub-question. A connected cluster answers a dozen, so it shows up across more of the fan-out. [Ahrefs](https://ahrefs.com/blog/topical-authority/) publishes a snapshot of the breadth this can produce: its page on magnesium glycinate at Healthline is tracked ranking for around 2,500 Google keywords and appearing in 473 AI Overview queries, 279 ChatGPT prompts, 200 Perplexity prompts, 28 Gemini prompts, and 86 Copilot prompts. Those are Ahrefs's tracked counts, so they show how wide a deeply covered page can reach, not proof of exactly why it was pulled each time. But the pattern is what depth looks like. One consequence changes how you write. Retrieval often fetches a single page without traversing your cluster, so **each page has to answer its own question on its own**. The cluster gives you more pages that can be retrieved; each retrieved page still has to stand alone once it is. That is where [content chunking](https://geotoolbox.ai/blog/content-chunking) matters, structuring each page so a self-contained passage can be lifted out and cited without the rest of the article around it. One thing the funnel hides: there is no single "AI search." Engines retrieve and cite differently, so authority on one does not transfer cleanly to the others. [Ahrefs found](https://ahrefs.com/blog/ai-search-overlap/) that only about 12% of the links cited by ChatGPT, Gemini, and Copilot also sit in Google's top 10 for the same prompt, while Google's own AI Overviews pull 76% from the top 10. Perplexity leans harder on freshness, engines weight off-site mentions differently, and a page ChatGPT reaches for may be one Gemini never surfaces. Build for the subject, but measure per engine. ## Is Topical Authority Even a Real Ranking Factor? Google has never published a formal specification for topical authority, and Ahrefs, whose own guide ranks at the top for the term, says so plainly. There is no documented, general topical authority score. The one place Google has confirmed a dedicated topic authority system is [news](https://developers.google.com/search/blog/2023/05/understanding-news-topic-authority), where it uses notability, influence, and source reputation to rank Top Stories. That is a specific, scoped system for news publishers, not a site-wide rating for the whole web. Patents show the wider industry has explored the idea directly; one held by Microsoft (US20210004416A1) describes producing a topical authority ranking for documents. But a patent describes something a company could do, not proof any engine runs it in live search. The research is honest about the limits too. A [2026 critical survey](https://arxiv.org/abs/2607.14035) that reviewed 45 studies on generative engine optimization found that while already-retrieved content can be nudged toward citation, no technique yet shows a stable, long-term, cross-platform causal effect on discoverability. Topical relevance and where a claim sits on the page were the strongest recurring levers, but nobody has proven a durable formula. So what is topical authority, really? It is a useful umbrella term for a set of things Google's systems actually do: read query intent rather than exact keywords (Hummingbird), model the relationships between concepts (BERT), and reward genuinely [helpful content](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) (a signal now folded into core ranking). [E-E-A-T](https://geotoolbox.ai/blog/eeat-ai-search) sits alongside these as the lens Google's quality raters use to judge expertise, not a score the algorithm reads directly. On the Google side, the most concrete lever is the one the 2024 API leak exposed: the tighter your content clusters around one subject, the cleaner the topical identity Google can build for your site. You are not optimizing for one hidden rating. You are optimizing for those underlying signals, which point you at what moves: coverage, structure, credibility, and being reachable. ## The Precondition Everyone Skips: Can AI Even Reach Your Cluster? All of the above assumes the engine can fetch and read your pages. Guides rarely make it the first step, and it is where clusters quietly fail. Retrieval starts with a crawler fetching the page, and the crawlers that feed AI answers are not always the ones people block for. The bot that lets ChatGPT cite you is OAI-SearchBot, not the GPTBot that only governs model training. Perplexity uses PerplexityBot. Google's AI Overviews are served from its search index by Googlebot, so blocking Google-Extended (which governs Gemini training) does nothing to remove you from them. Block the wrong bot, or serve content that only appears after JavaScript the crawler does not run, and the engine you built the cluster for never sees it. Which crawlers matter depends on which engines you care about, so it is worth knowing [how each AI crawler behaves](https://geotoolbox.ai/blog/ai-crawlers) rather than blocking or allowing them by reflex. The common ordering mistake is to plan thirty new articles while the dozen you already have return an empty page to a search crawler. Coverage you cannot serve is coverage that does not count. The fix is cheap and comes first: confirm the pages you already have are reachable and rendered before you write more. Our [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) reports, crawler by crawler, whether each of the AI bots gets a real page or a wall. It is a five-minute check that decides whether the next three months of content work can be seen at all. Reachability is not a growth tactic; it is the floor everything above it stands on. ## How to Build Topical Authority The work is a sequence, not a checklist you tackle in any order. Each step feeds the next. 1. **Map the topic first.** Start from a seed subject specific enough to own ("coffee roasting," not "coffee"), then expand it into the subtopics, questions, and intents a real expert would cover. Decide what to leave out with the same care you decide what to include; a map that wanders dilutes focus. 2. **Build it as a topic cluster.** This is where the topic cluster earns its keep. Publish one pillar page that covers the subject broadly and acts as the hub, then supporting cluster pages that go deep on each subtopic. Link every cluster page back to the pillar and the pillar out to each cluster, and cross-link cluster pages that are genuinely related. That internal structure is what tells an engine the pages are one connected body of work rather than scattered posts. The pillar-and-cluster model was popularized by [HubSpot](https://blog.hubspot.com/marketing/topic-clusters-seo), and it remains the cleanest way to make coverage legible. If a subtopic is too thin to stand alone, fold it into a neighbor instead of publishing a weak page. 3. **Write for depth and for extraction.** Match the search intent of each page, cover its subtopic thoroughly, and show real expertise rather than a rewrite of the top results. Then make it extractable: lead with the answer, use clear headings, and include the specifics engines lift. The foundational [GEO study](https://arxiv.org/abs/2311.09735) found that adding statistics, quotations, and cited sources can raise a page's visibility in generative answers by up to 40%. Concrete, sourced writing is both better for readers and more citable. 4. **Earn topically relevant signals off-site.** Links and mentions from sites inside your subject reinforce the association more effectively than unrelated ones, even strong ones, though link quality and prominence still matter. Relevance is the part specific to topical authority, which is one reason [where you build links](https://geotoolbox.ai/blog/best-link-building-platforms) matters as much as how many. 5. **Stay on-topic and prune what strays.** Google's 2024 API leak exposed two internal fields, a site focus score for how concentrated your content is and a site radius for how far it wanders. Their exact weight in live ranking is not documented, but the implication fits everything else here: content far outside your core subject is more likely to blur your topical identity than add to it. When old pages have pulled your radius wide, consolidate, redirect, or remove them rather than leaving them to muddy what your site is about. One structural nicety: you can describe the pillar-to-cluster relationship with [schema markup](https://geotoolbox.ai/blog/schema-markup-for-ai), declaring which pages are parts of a larger whole. Engines may or may not use it for topic clustering, and it will not manufacture authority, but it makes your structure explicit rather than merely implied. ## How to Measure Topical Authority When There Is No Score Start from the uncomfortable premise: there is no native topical authority score to read off a dashboard. The "Low / Medium / High" labels some tools show are proprietary estimates, not a Google number. So you measure by proxy, and the proxies below are the ones that track the outcome.
ProxyWhat it tells youWhere to see it
Branded topic-query varietyYour audience associates your brand with the subjectSearch Console (queries mixing brand + topic)
Unbranded cluster rankingsCoverage is compounding across the subject, not one pageRank tracking / share of voice for the cluster
Topic traffic shareWhether you are becoming the destination for the subjectAnalytics + competitor share estimates
AI share of voice and citationsEngines are pulling you into answersAn AI visibility tracker across engines
One cheap early proxy from practitioners: count how many of your subject's keywords sit in positions 9 to 20 in Search Console. A cluster that is being recognized tends to push a growing band of terms to the edge of page one before any of them breaks through, so that band widening is an early sign the coverage is landing. The most modern proxy, and the one most relevant if AI search is the reason you care about topical authority, is [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice): how often you get mentioned or cited across ChatGPT, Perplexity, Gemini, and AI Overviews for the prompts in your subject. It is measurable now, and unlike a keyword ranking it tells you whether you are getting retrieved and cited at all. Pair it with an [AI visibility score](https://geotoolbox.ai/blog/ai-visibility-score) to watch the trend rather than a single snapshot. None of these is a single number that says "you have arrived." Read together over time, they tell you whether your coverage is being recognized, which is the only thing topical authority was ever really claiming to be. ## Where Topical Authority Backfires The concept has failure modes, and they are common enough to name. **Thin, scaled clusters.** Publishing 300 shallow, largely AI-written pages to blanket a topic does not build authority. Enough low-value content can drag down how Google's core ranking systems treat the whole domain. Thirty genuinely deep, interconnected pages beat 300 thin ones for both ranking and citation. **Chasing volume off-topic.** Every article outside your subject widens your site radius and blurs what you are about. Traffic potential is not a good enough reason to cover something your site has no credible claim to. **Coverage no one can reach.** The quietest failure: a complete cluster behind a crawler block, or one that only renders in JavaScript, earns nothing. **Treating it as finished.** Topical authority decays. Outdated stats, superseded tools, and stale pages drop you out of AI Overviews you used to appear in. Refresh on real triggers, a traffic drop, a superseded source, or a lost citation, rather than on a calendar. ## Building Authority You Can Actually Prove Topical authority is not a score you switch on. It is what search and AI engines conclude once you have covered a subject deeply, connected the coverage so the relationships are legible, backed it with real expertise, and, the part most teams skip, made sure the engines can reach and read it. That last point is where we spend most of our time. Before you commit a quarter to a cluster, it is worth confirming the pages you already have can be fetched and rendered by the crawlers that matter. Our [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) reports which AI crawlers each page lets through, and the broader [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) flags the reachability and structure problems that keep well-written clusters out of AI answers. Cover the subject, then make sure it can be seen. ## Frequently Asked Questions ### How long does it take to build topical authority? Ahrefs's guidance puts it at six to twelve months before meaningful movement, and that matches most practitioner experience: coverage first, early ranking gains in the middle, compounding later. Treat it as a rough heuristic, not a schedule. It varies with your competition, your starting authority, and how fast you publish. ### How many articles do I need? There is no fixed number. A tightly defined niche might need 15 to 20 well-connected pieces; a broad subject can need 50 or more. The better question than "how many" is whether your coverage answers the subject more completely than the sites you are competing with. ### Is topical authority the same as domain authority? No. Domain authority (or Ahrefs' Domain Rating) estimates the backlink strength of your whole site and is not topic-specific. Topical authority is subject-level depth in one area. That gap is why a small, focused site can win the subject it owns against a much stronger generalist. ### Does topical authority guarantee AI citations? No. It improves the odds by getting your pages into the retrieval pool, but the engine still has to reach the page, find a self-contained passage that answers the sub-question, and prefer it over alternatives. Reachability, extractable structure, and topically relevant off-site signals all have to be in place too. ### Does AI-generated content build topical authority? Only if it is genuinely useful and reviewed. Depth and accuracy build authority regardless of how a draft started. Thin, unedited AI content published at scale does the opposite: enough of it can weigh on how core ranking systems treat the whole site. ## Sources - Topical Authority: What It Is, How Google Measures It, and How to Build It - Ahrefs, Despina Gavoyannis - `ahrefs.com/blog/topical-authority/` - Understanding news topic authority - Google Search Central, 2023 - `developers.google.com/search/blog/2023/05/understanding-news-topic-authority` - Creating helpful, reliable, people-first content - Google Search Central - `developers.google.com/search/docs/fundamentals/creating-helpful-content` - AI Search Overlap: How Often AI Citations Match Google's Top Results - Ahrefs - `ahrefs.com/blog/ai-search-overlap/` - Topical Authority (Holistic SEO) - Koray Tuğberk Gübür - `holisticseo.digital/theoretical-seo/topical-authority/` - GEO: Generative Engine Optimization - Aggarwal et al., arXiv 2311.09735 - `arxiv.org/abs/2311.09735` - Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026) - arXiv 2607.14035, July 2026 - `arxiv.org/abs/2607.14035` - Topic Clusters: The Next Evolution of SEO - HubSpot - `blog.hubspot.com/marketing/topic-clusters-seo` --- ## What Is Meta AI? The Assistant, the Muse Models, and Why It Matters > What is Meta AI? A current guide to Meta's free assistant: the Llama-to-Muse Spark switch, where it runs, Meta AI vs ChatGPT, privacy, and AI visibility. - Canonical: https://geotoolbox.ai/blog/what-is-meta-ai - Published: 2026-08-16 · Updated: 2026-08-21 Meta AI is the assistant that showed up uninvited in your WhatsApp, your Instagram search bar, and that blue-and-pink circle you keep tapping by accident. It answers questions, makes images, and talks back through Ray-Ban glasses, and more than a billion people use it every month, most of them without ever choosing to. If you want the plain version of what Meta AI actually is, what powers it now, where it lives, and whether you can trust it, this is it, current as of August 2026. Most "what is Meta AI" articles tell you it runs on Llama. As of 2026, that is out of date, and the gap matters. In April 2026 Meta quietly moved its flagship assistant onto a brand-new proprietary model called Muse Spark, its first model built outside the open Llama line. We will cover what Meta AI is, the switch from Llama to Muse, everywhere it runs, how it compares to ChatGPT and Gemini, the privacy questions worth asking, and one thing almost no marketer is tracking: whether your own site shows up when Meta AI answers.
![Timeline of Meta's AI models from FAIR in 2013 through Llama to the Muse Spark, Muse Image and Muse Video models in 2026.](/blog/what-is-meta-ai/meta-ai-model-timeline.png)
Meta AI's model line moved from open-weight Llama to the proprietary Muse family in 2026.
## What Is Meta AI? **Meta AI is the free artificial intelligence assistant built by Meta, the company behind Facebook, Instagram, WhatsApp, and Messenger.** It is a general-purpose chatbot in the same category as ChatGPT, [Google Gemini](https://geotoolbox.ai/blog/what-is-gemini), and Claude: you ask it questions in plain language and it answers, writes and edits text, generates and edits images, holds voice conversations, and can search the web for current information. What makes it different is not the technology but the distribution. Instead of asking you to visit a separate site, Meta built the assistant directly into apps that billions of people already open every day. One naming point trips people up. The label "Meta AI" gets used for more than one thing. It is the name of the consumer assistant this article is about, and it is also the older name for Meta's research division, formerly Facebook AI Research (FAIR), founded in 2013, which builds the underlying models. When Wikipedia says "Meta AI is a research division," it means the lab. When the blue circle in WhatsApp says "Meta AI," it means the assistant. This article is about the assistant, and the [large language model](https://geotoolbox.ai/glossary/large-language-model) that runs it. The assistant is genuinely free. There is no subscription to chat with it, generate images, or use it in the apps. You pay in a different currency, which is your data, and we will get to exactly how in the privacy section below. ## Which Model Powers Meta AI? From Llama to Muse Spark **Meta AI used to run on Llama, Meta's family of open-weight models. As of April 2026, the flagship Meta AI app and the meta.ai website run on Muse Spark, a new proprietary model, and this is the single most out-of-date fact in most explainers.** If you have read that "Meta AI is powered by Llama 4," that was true through early 2026 and is no longer the whole story. Here is what happened, and the "why" explains the whole shift. For years Meta's strategy was open source: it released Llama 2, Llama 3, and Llama 4 as downloadable models anyone could run, which made Llama one of the most-used open model lines in the world. Then Llama 4 stumbled. According to [CNBC's reporting](https://www.cnbc.com/2026/04/08/meta-debuts-first-major-ai-model-since-14-billion-deal-to-bring-in-alexandr-wang.html), the largest planned model, Llama 4 Behemoth, was delayed through 2025 and then frozen, and the disappointment triggered a leadership shake-up. Meta paid roughly $14.3 billion for a 49% stake in Scale AI and hired its chief executive, Alexandr Wang, as Meta's Chief AI Officer to run a new group called Meta Superintelligence Labs, with Shengjia Zhao as chief scientist. The first product of that group was Muse Spark, released April 8, 2026. Meta describes it as a natively multimodal reasoning model with tool use and multi-agent capabilities, and says it matches its previous model, Llama 4 Maverick, while using [over an order of magnitude less compute](https://ai.meta.com/blog/introducing-muse-spark-msl/). Reporters framed the launch as "goodbye Llama," but Meta itself is more careful: its announcement says Muse Spark is "available today at meta.ai and the Meta AI app," and does not claim it has replaced Llama across WhatsApp, Instagram, and Facebook, where the embedded assistants have historically run on Llama. So the accurate statement in 2026 is narrower than the headlines: the standalone Meta AI app and meta.ai run on Muse Spark, while the assistant inside the social apps is mid-transition. Muse Spark is also the anchor of a wider family of models Meta shipped through 2026.
ModelReleasedWhat it does
Muse SparkApril 8, 2026The reasoning model that powers the Meta AI app and meta.ai; multimodal, tool use, agentic tasks
Muse Image + Muse VideoJuly 7, 2026Meta Superintelligence Labs' first media-generation models; Muse Image launched in the Meta AI app and on meta.ai, plus Instagram Stories in the US and WhatsApp in limited countries; Muse Video shipped as a preview
Muse Spark 1.2 + Muse CodeEarly August 2026Muse Spark 1.2, a coding-focused model (an update to Muse Spark 1.1), powering Muse Code, a terminal-based coding agent (beta), and the Meta Model API; it does not power the meta.ai assistant
Llama 4 (Scout, Maverick)April 2025The prior open-weight line; still available to developers and still powering some in-app assistants
One nuance keeps the "Meta went closed" reading honest: on August 10, 2026 Meta released Muse Glimmer, a 30-billion-parameter multimodal model with open weights under an Apache 2.0 license on Hugging Face, meant for local and on-device use. So Meta now runs two tracks in parallel: a closed frontier model (Muse Spark) and open weights anyone can run themselves (Glimmer). ## Where You Can Use Meta AI Meta AI's whole strategy is to be everywhere you already are, which is why it seems to appear on your phone without your permission. You do not download Meta AI so much as find it waiting inside apps you installed for other reasons. That is also the answer to the most common question about it, which is "why did this show up on my phone?" It showed up because Meta added it to the app, not because you did anything. You can reach the assistant on the web, in its own app, inside Meta's social apps, and on its hardware. On the web there is **meta.ai**, which works in any desktop or mobile browser, alongside the standalone **Meta AI app** for iOS and Android; these are the surfaces now running on Muse Spark. It is built into **WhatsApp** (the blue-and-pink circle), **Instagram** (the search bar and direct messages), **Facebook**, and **Messenger**. And it reaches hardware through the **Ray-Ban Meta smart glasses**, where you talk to it out loud and it can describe what the camera sees. The experience varies by surface: the app and website get the newest features first, the in-app versions are more limited, and availability differs by country. ## Meta AI vs ChatGPT vs Gemini: Is It Any Good? Meta AI is good enough for everyday questions and casual image generation, and its real advantage is convenience rather than raw capability. Because it is already inside the apps you use, it wins on friction: you never have to open anything new. On hard reasoning, long research tasks, and coding, the more focused assistants tend to lead, and Meta says so itself. Its own Muse Spark announcement lists [long-horizon agentic systems and coding workflows](https://ai.meta.com/blog/introducing-muse-spark-msl/) as areas with current performance gaps. Users also report the two rough edges you would expect from a free, mass-market assistant: it can state wrong facts confidently, and its tone runs noticeably chummier than its rivals.
AssistantUnderlying modelBest forCostWatch for
Meta AIMuse Spark (app + meta.ai)Quick answers and images inside apps you already useFreeConfident errors; data used for ads
ChatGPTGPT-5.xReasoning, writing, coding, deep researchFree tier + paidPaid tiers for best models
GeminiGemini 3.xGoogle-ecosystem tasks, long context, multimodalFree tier + paidTied to Google account and services
So is Meta AI better than ChatGPT? For most people, the honest answer is that it is the most *convenient*, not the most capable. If you want a quick answer while messaging a friend or a fast image in a group chat, Meta AI is right there and free. If you are doing serious reasoning, research, or code, ChatGPT and [Gemini](https://geotoolbox.ai/blog/gemini-vs-chatgpt) are still the stronger tools. Many people end up using both, which is exactly what Meta is counting on to keep the assistant in front of a billion users. ## Can Your Site Show Up in Meta AI? **Yes, and it is the AI-search surface most marketers overlook.** When Meta AI answers a question that needs current information, it searches the web, and Meta runs its own crawler, [Meta-ExternalAgent](https://geotoolbox.ai/glossary/meta-externalagent), to index public pages that can be cited or linked in those answers. That means your website is a candidate to appear inside Meta AI responses the same way it can appear in ChatGPT search or Google's AI Overviews. The levers are the same ones that drive [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) everywhere: the crawler has to be able to reach your pages, your content has to answer the question cleanly enough to be lifted out, and your facts have to be current. The scale is what makes ignoring it a mistake. Meta AI crossed a billion monthly active users in May 2025, per [Mark Zuckerberg at Meta's shareholder meeting](https://www.cnbc.com/2025/05/28/zuckerberg-meta-ai-one-billion-monthly-users.html), and it has kept growing since. That is a larger audience than most AI assistants marketers actually optimize for, yet nearly all the attention goes to ChatGPT and Gemini. There is a second reason this matters, and it is the reason this article exists. AI assistants answer from their training data unless they search, and training data goes stale. When we ran a quick test on August 16, 2026, asking a current version of ChatGPT with web search turned off what powers Meta AI, it confidently answered "the Llama family, currently chiefly Llama 4," months after Muse Spark had shipped. In the AI-visibility work we do, that is the pattern that decides who gets cited: engines fall back on outdated training data, and the pages that get pulled into answers are the current, clearly structured, reachable ones. Being the current source on your own topic is the mechanism, not a bonus. ## Privacy and Control: Turning Meta AI Off and What It Does With Your Data Meta AI is wired into Meta's data machine, so the sensible approach is to use it for low-stakes things and keep private details out of it. Start with the question most people actually have: no, Meta AI does not read your normal WhatsApp messages. One-to-one and group chats stay end-to-end encrypted, and the assistant only sees a message when you deliberately tag @Meta AI or message it directly. The catch is that a message you do send to Meta AI is not private to Meta, and it can be used the way the rest of your AI activity is. Turning it off is harder than hiding it. WhatsApp is rolling out a "Show Meta AI Button" toggle under Settings that removes the blue circle, but it is phased, is not yet available in the US and much of Asia, and even where it exists it only hides the button rather than uninstalling the feature. Everywhere else you are limited to muting or archiving the chat. Available controls also depend on your country: Meta AI arrived later, and with more restrictions, in the EU than in the US. On data, a few things are worth knowing. Meta uses public posts, photos, and captions from adult Facebook and Instagram accounts to train its models, and a formal opt-out is largely limited to the EU, UK, and a few other regions. As of December 16, 2025, Meta also began [using your AI interactions to shape the content and ads](https://about.fb.com/news/2025/10/improving-your-recommendations-apps-ai-meta/) you see in most regions (the [EU, UK, and South Korea are excluded](https://www.forbes.com/sites/kateoflahertyuk/2025/10/06/meta-ai-confirms-your-data-will-be-used-for-ads-heres-how-and-when/) under GDPR and similar laws), with some exceptions for sensitive topics. And in 2025, some Meta AI app users published chats to a public "Discover" feed by tapping share without realizing the posts were public, a reminder to treat any chatbot like a semi-public space. The practical rule is the one that applies to any AI assistant: do not paste passwords, financial details, medical information, or anything you would not want tied to your identity. And if you wear the Ray-Ban glasses, the camera captures the people around you too, which is why some venues have started to restrict them. ## The Short Version Meta AI went from an open-Llama experiment to a proprietary, Muse-powered product with more than a billion monthly users, embedded so deeply in Meta's apps that most people never chose to use it. It is free and convenient, uneven on hard tasks, and quietly busy with your data. For anyone who runs a website, though, the headline is simpler: Meta AI is a search surface with a billion users and its own crawler, and hardly anyone is measuring whether they show up in it. That last part is worth thirty seconds of your time. You can check whether Meta's crawler can even reach your site with our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker). If Meta-ExternalAgent cannot read your content, you are invisible to a billion-user assistant before the ranking question even starts. ## Frequently Asked Questions ### Is Meta AI free? Yes. Chatting with Meta AI, generating images, and using it inside WhatsApp, Instagram, Facebook, and Messenger are all free, with no subscription. The trade-off is data: Meta uses public posts to train its models and, as of late 2025, uses your AI interactions to personalize ads in most regions (the EU, UK, and South Korea are excluded). ### What model does Meta AI use now? As of August 2026, the Meta AI app and the meta.ai website run on Muse Spark, a proprietary model released in April 2026 by Meta Superintelligence Labs. This replaced Llama on those surfaces. The assistants built into WhatsApp, Instagram, and Facebook have historically run on Llama and are mid-transition. ### Can I turn off Meta AI? There is no universal off switch. WhatsApp is rolling out a "Show Meta AI Button" toggle that hides the blue circle, and in most apps you can hide, mute, or archive the Meta AI chat, but you cannot fully remove the feature. Data controls also vary by region: users in the EU and UK generally have stronger opt-out and data rights than those elsewhere. ### Is Meta AI safe to use? It is safe enough for everyday questions, but treat it like any public-facing chatbot. It can state wrong facts confidently, and Meta uses conversation and account data to personalize ads in most regions (the EU, UK, and South Korea are excluded). Do not share passwords, financial details, medical information, or anything sensitive. ### Is Meta AI better than ChatGPT? For convenience, often yes, because it is already inside the apps you use and it is free. For hard reasoning, research, and coding, ChatGPT and Gemini are still stronger. Many people use Meta AI for quick tasks and a dedicated assistant for serious work. ### Who makes Meta AI? Meta AI is built by Meta, the parent company of Facebook and Instagram, through its research arm and its newer Meta Superintelligence Labs. The assistant is simply called "Meta AI," and the models under it are the Llama and, more recently, Muse families. ## Sources - Introducing Muse Spark - AI at Meta, April 8, 2026 - `ai.meta.com/blog/introducing-muse-spark-msl` - Meta debuts first major AI model since $14 billion deal to bring in Alexandr Wang - CNBC, April 8, 2026 - `cnbc.com/2026/04/08/meta-debuts-first-major-ai-model-since-14-billion-deal-to-bring-in-alexandr-wang.html` - Mark Zuckerberg says Meta AI has 1 billion monthly active users - CNBC, May 28, 2025 - `cnbc.com/2025/05/28/zuckerberg-meta-ai-one-billion-monthly-users.html` - Introducing Muse Image and Muse Video - AI at Meta, July 7, 2026 - `ai.meta.com/blog/introducing-muse-image-muse-video-msl` - Meta AI - Wikipedia - `en.wikipedia.org/wiki/Meta_AI` - Meta Superintelligence Labs - Wikipedia - `en.wikipedia.org/wiki/Meta_Superintelligence_Labs` --- ## Zero-Click Searches: The Real Numbers and What Still Works in 2026 > Zero-click searches now end ~68% of US Google queries. Here is what that number really means, why rankings hold while traffic falls, and the levers that still work. - Canonical: https://geotoolbox.ai/blog/zero-click-searches - Published: 2026-08-16 · Updated: 2026-08-16 Your rankings are steady and your traffic is sliding anyway. That gap has a name: zero-click searches, the queries that now end on the results page without anyone clicking through. In the US, about 68% of Google searches end that way, and AI Overviews are pushing the number higher every year. This is the honest version. What zero-click actually is, how big it really is once you sort out the competing numbers, why ranking first no longer guarantees the click, and the specific levers that still work when people stop clicking.
![Bar chart of the US Google zero-click rate: 58.5% and 60.45% in 2024 (both on the Datos panel) rising to 68.01% in early 2026 (Similarweb panel).](/blog/zero-click-searches/zero-click-trajectory.png)
The 2024 readings use SparkToro's Datos panel; the 2026 reading switched to Similarweb, so the jump is partly a change of measurement. Source: SparkToro 2024 and 2026 zero-click studies.
## What Is a Zero-Click Search? A zero-click search is a query that ends on the results page itself, with no click through to any website. The searcher reads the answer, then closes the tab or types a new query. You can see the full definition in our [zero-click search glossary entry](https://geotoolbox.ai/glossary/zero-click-search), but the mechanics matter more than the label. Zero-click happens because the results page increasingly answers the question for you. The features doing the answering are familiar: AI Overviews at the top, featured snippets, knowledge panels, People Also Ask boxes, local packs, and instant widgets like the weather, a calculator, or a sports score. Each one resolves a certain class of query without sending anyone to a site. One thing to be clear about early: zero-click is not a metric Google reports. You will not find it in Search Console. It is a behavioral outcome measured by clickstream panels, where a research firm watches a sample of real users and reconstructs what happened after each search: a click to the open web, a click to a Google property, or no click at all. That distinction shapes every number in the next section. ## How Big Is Zero-Click, Really? (And Why the Numbers Disagree) The headline figure is **68% of US Google searches ended without a click** in the first four months of 2026, from [SparkToro's 2026 study](https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/) built on Similarweb clickstream data. That number gets quoted as 58.5%, 60.45%, and 68% across different articles, and the gap is not sloppiness. It is what happens when the measuring instrument changes underneath the trend. Here are the figures SparkToro actually reports, with the panel each one came from.
FigureSourcePanelWindowWhat it measures
68.01% USSparkToro 2026Similarweb clickstreamJan–Apr 2026US searches that ended with no click
60.45% USSparkToro 2026 (2024 baseline)Datos clickstream2024The 2024 figure the 2026 study compares against
58.5% US / 59.7% EUSparkToro 2024 studyDatos clickstreamSep 2022 – May 2024The earlier study, same panel, different window
Here is the trap. SparkToro measured 2024 with Datos and 2026 with Similarweb, different clickstream panels with different users and devices. SparkToro says so itself, calling the comparison "a bit of apples and oranges." So the honest read is that zero-click is clearly climbing, but the exact size of the jump is fuzzy, because the ruler changed between readings. Both 2024 figures, 58.5% and 60.45%, come from Datos over slightly different windows, which is the whole reason you should never lift a number from one study, set it beside a number from another, and call the difference a trend. What is not in doubt is the direction. AI Overviews are the clearest accelerant: [Ahrefs data](https://ahrefs.com/blog/ai-overview-triggers/) that SparkToro cites puts them on more than 20% of searches, and when one appears, [click-through to the results below drops by close to 60%](https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/). For a longer read on where these numbers come from, see our breakdown of the [state of AI search in 2026](https://geotoolbox.ai/blog/state-of-ai-search-2026). One footnote worth keeping straight: the phrase "zero-click search" itself gets searched a little less than it did a year ago, even as the phenomenon behind it hits record highs. The term cooling is not the problem cooling. ## Why Your Rankings Hold but Your Traffic Falls If you have watched a page sit at position one while its clicks slide toward zero, you are not imagining it. The rank did not move. The click did. The reason is simple once you name it. Ranking first used to mean you got the click, because the answer lived on your page. Now the answer often lives on the results page. When an AI Overview or a featured snippet resolves the query, the classic blue links below it become optional, and most people never scroll. That is why click-through to those links falls so sharply, even for the pages ranking directly beneath the answer. We cover how those summaries get built in [what Google AI Overviews are](https://geotoolbox.ai/blog/what-are-google-ai-overviews). This shows up in Search Console as a specific, recognizable pattern: impressions flat or rising, clicks falling, average position steady. Your content is being shown and read. It is just being read inside Google's answer instead of on your site. Ranking number one now can mean almost no clicks at all, and the instinct is to treat it as a penalty or an algorithm hit. It usually is neither. The unhelpful conclusion is that you should rank harder. You cannot out-rank a layout change. When the answer is delivered above your listing, a higher position does not buy back the click. That reframes the whole problem: the job is no longer only to rank, it is to be the answer and to capture the intent that still converts. The rest of this guide is about how. ## Cited but Not Clicked: What Visibility Is Actually Worth Here is the part most guides skip. Being named in an answer is not the same as being visited, and the gap is enormous. In [Pew Research Center's 2025 study](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) of 900 US adults across 68,879 searches, on the visits where an AI summary appeared, users clicked one of its cited source links in about 1% of cases. A citation is real estate in the answer, not a referral to your site. So is that visibility worth anything? The answer depends on who you ask, and both camps have a point. The skeptical read, argued well by Search Engine Land, is that a citation nobody clicks is brand awareness at best and has no direct business impact. If the reader never reaches your page, your calls to action and your funnel never load. The optimistic read, from Semrush and others, is that repeated presence in answers builds trust and pulls people back later through branded search or a direct visit. Both can be true, and where you land depends on the query. For a definitional question the reader was never going to convert on, a citation is mostly a branding impression. For a commercial question, being the cited source shapes who the buyer trusts before they ever compare options. There is a thin upside hiding in that 1%, though: the people who do click through have already read the gist, so they arrive better informed than a cold searcher would. The defensible middle is this: treat citations as a real signal, but measure them as demand, not as traffic, and expect the value to show up downstream rather than in the same session. ## Zero-Click Isn't Just a Google Problem Most coverage frames zero-click as a Google story. It is bigger than that. When someone asks ChatGPT, Perplexity, Gemini, or an AI browser a question, there is no results page at all. The model answers directly, and if your brand is not in that answer, you are invisible in a way that has no ranking to check and no snippet to win. That is zero-click taken to its limit: not a click lost from a SERP, but a SERP that never appears. This matters for measurement more than for panic. If you only watch Google, you are watching one surface of a problem that now spans several. A brand can be summarized accurately by Google, misquoted by one assistant, and absent from another, all at the same time, and none of it shows up in a rank tracker. Seeing the whole picture means tracking how often you are mentioned and cited across engines, not just where you rank on one. Cross-engine mention tracking is what our tools at geotoolbox are built for. If you want the mechanics, we walk through [how to track brand mentions in AI search](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search). ## How to Measure Zero-Click Impact on Your Own Site You can measure this with tools you already have, before paying for anything. 1. **Find the erosion in Search Console.** Filter to the last six to twelve months and look for pages where impressions held or rose while clicks and CTR fell. That divergence is the signature of zero-click: you are being shown and read, just not visited. One caveat, straight from [Google's own documentation](https://developers.google.com/search/docs/appearance/ai-features): AI Overview appearances are folded into your normal Search Console traffic under the "Web" type. There is no dedicated filter to isolate them, so treat the impressions-up-clicks-down pattern as your best available proxy, not a clean read. 2. **Watch branded search and direct traffic.** When people meet your brand inside an answer and do not click, the influence often surfaces later as a branded query or a direct visit. Rising branded impressions in Search Console are a reasonable proxy that your presence in those answers is registering. 3. **Read the AI Assistant channel in GA4.** In 2026 Google [added a native AI Assistant channel to GA4](https://www.searchenginejournal.com/google-analytics-adds-ai-assistant-as-default-channel-group/574974/) that recognizes referrals from assistants like ChatGPT, Gemini, and Claude. A couple of limits keep it from being a full count. AI Overview clicks are not in it at all, because Google reports those as ordinary organic search. And any assistant Google has not added to its recognized list (Perplexity, at the time of writing), plus the many AI answers that arrive with no referrer, still land in Referral or Direct. Treat it as a floor on AI-driven traffic, not a full count, and lean on steps 1, 2, and 4 for the rest. 4. **Track citations across engines, not just clicks.** The metric that most directly tracks whether you are present is how often you are mentioned and cited in AI answers, measured over time. That is a visibility number, not a sessions number, and it is the one closest to the outcome you now care about. Our guide on [how to measure your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) lays out the full approach. None of this requires new schema or a rebuild. It requires reading the tools you have for a different signal. ## What Actually Still Works The fatalist take is that zero-click means SEO is dead and there is nothing you can do. That is wrong, and the people saying it are usually treating every query the same. The shift is not from search to nothing. It is from capturing traffic to capturing demand and citations. Here is what that looks like in practice. **Start by triaging your queries.** Some still earn the click and some never will again. Spend your effort where the click survives, and optimize the rest to be the cited answer rather than fighting a layout you cannot win.
Query typeStill earns the click?What to do
Definitional, "what is", quick factsRarelyConcede the click; structure the passage to be the cited source
Generic how-toDecreasinglyAnswer-first for citation; do not over-invest
High-intent commercial, pricing, "best X for Y"YesFight for it: sharp titles, a page that converts the visit
Comparison, "X vs Y", buying decisionsYesGo deep with original data a summary cannot compress
Branded and navigationalYesOwn it; this is where zero-click demand lands later
**Become the source worth citing.** For the queries you concede, the goal is to be the passage the answer is built from. Lead sections with a direct, self-contained answer, then support it with something a model cannot flatten: proprietary data, a real framework, first-hand results. If your content can be summarized without losing anything, it will be, and you will not get credited. Our guide on [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) goes deeper on the formatting. **Capture demand off the click.** Branded search, email capture, and community are how you stay reachable when Google and the assistants keep the session. Treat every zero-click impression as the top of a funnel that ends somewhere you own. A couple of myths to drop. Blocking AI crawlers will not save your Google traffic: AI Overviews are built from Google's own index, so short of removing a page from search entirely with noindex, you cannot opt a ranked page out of them. Blocking an assistant's crawler is a real lever but a blunt one. It can lower your odds of being cited by that engine, though the bots differ (OpenAI alone runs one crawler for training and another for search) and some assistants can still surface you from third-party pages. Mostly it costs you AI visibility without buying back a single Google click. And there is no special "AI schema" to chase. Google states plainly that no structured data is required to appear in AI Overviews or AI Mode. Standard schema that matches your visible content is still worth having, but not as an AI shortcut. If you are still asking whether any of this means the end of the discipline, we answered that directly in [is SEO dead](https://geotoolbox.ai/blog/is-seo-dead). ## The Shift Is Real, and So Is the Response Zero-click is the search page doing more of the answering, and the trend line points one way. Waiting it out is not a plan. You will not win the old clicks back by ranking harder. What you can do is decide which queries still deserve the fight, become the source the answers are built from, and measure whether your presence in those answers is actually growing. That last part is where most teams are flying blind. If you want to see where your brand is being cited and mentioned across Google's AI Overviews, ChatGPT, Perplexity, and the rest, [geotoolbox tracks your AI visibility across engines](https://geotoolbox.ai/features/domain-overview) so you can watch the metric the click no longer captures on its own. The traffic you lost to zero-click is not coming back in its old form. Being the answer, and knowing when you are, is the version that does. ## Frequently Asked Questions ### Is it true that 60% of searches are zero-click? It is higher than that now. The most current figure is about 68% of US Google searches ending without a click in early 2026, per SparkToro's Similarweb-based study. The 58.5% and 60.45% figures are 2024 numbers, both from the Datos panel, while the 68% comes from a different panel (Similarweb), so the studies are not directly comparable. ### How many searches are zero-click in 2026? Roughly 68% in the US, per SparkToro's 2026 Similarweb data. It is one of the fastest-rising numbers in search, up from a Datos-panel reading in the low 60s in 2024, and AI Overviews are the clearest driver. ### Are zero-click searches killing SEO? They are killing one version of it: rank first, collect the click. The discipline itself is moving from winning traffic to earning demand and citations. High-intent, commercial, comparison, and branded queries still earn clicks. Generic definitional and how-to queries increasingly do not, and for those the goal becomes being the cited source. ### Should I block AI crawlers to stop zero-click searches? No. Blocking AI crawlers will not stop Google's own AI Overview from summarizing a page you already rank for, and blocking an assistant's crawler mostly just lowers your odds of being cited there. You lose AI visibility without recovering the Google click. ### How do I measure zero-click search impact? Look in Search Console for pages where impressions held or rose while clicks and CTR fell, which is the zero-click signature. Track branded search and direct traffic as downstream signals, use GA4's AI Assistant channel for ChatGPT, Gemini, and Claude referrals, and monitor how often you are cited across AI engines rather than only how you rank. ### Do zero-click citations actually convert? Not directly and not often in the same session. Pew found that on visits where an AI summary appeared, users clicked one of its cited links in about 1% of cases. The value shows up downstream, as branded search and better-informed visits later, and it is stronger for commercial queries than for purely informational ones. ## Sources - SparkToro - In 2026, Less than One Third of Google Searches Still Send a Click, 2026 - `sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click` - SparkToro - 2024 Zero-Click Search Study, 2024 - `sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches-only-374-clicks-go-to-the-open-web-in-the-eu-its-360` - Ahrefs - AI Overview triggers and their effect on click-through rate, 2025 - `ahrefs.com/blog/ai-overview-triggers` and `ahrefs.com/blog/ai-overviews-reduce-clicks-update` - Pew Research Center - Google users are less likely to click on links when an AI summary appears, July 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` - Google Search Central - AI Features and Your Website, 2026 - `developers.google.com/search/docs/appearance/ai-features` - Search Engine Journal - Google Analytics Adds AI Assistant as Default Channel Group, 2026 - `searchenginejournal.com/google-analytics-adds-ai-assistant-as-default-channel-group/574974` --- ## ChatGPT Ads: How They Work, What They Cost, and Who Sees Them > ChatGPT ads are live for Free and Go users. How they work, who sees them, what advertisers pay, and what they mean for your organic AI visibility. - Canonical: https://geotoolbox.ai/blog/chatgpt-ads - Published: 2026-08-15 · Updated: 2026-08-17 ChatGPT has ads now. A US pilot began in February 2026, and OpenAI has built it into a full ad platform since: Free and Go users, logged in or out, now see a labeled "Sponsored" card below some answers. If you are a marketer, you want to know how ChatGPT ads work, who they reach, and what they cost. If you are a regular user, you probably want to know why ads showed up and whether they change the answers. This covers both, plus the question most guides skip: what paid ads mean if what you actually want is for ChatGPT to recommend your brand on its own. ## Does ChatGPT Have Ads Now? Yes. OpenAI [began testing ads in ChatGPT on February 9, 2026](https://techcrunch.com/2026/02/09/chatgpt-rolls-out-ads/), starting in the United States, after laying out its advertising approach in a [January 16 post](https://openai.com/index/our-approach-to-advertising-and-expanding-access/). What launched as a small US pilot has since expanded to more countries and grown into a self-serve platform that advertisers can buy directly. ChatGPT is not the first AI answer engine to carry ads, Perplexity has run sponsored questions and Google's AI Overviews show ads too, but it is the largest to flip the switch. The important framing, straight from OpenAI: an ad is a paid placement, not a recommendation from ChatGPT. The model's answer and the ad beneath it come from separate systems. That distinction is small until you are the one trying to show up in the answer, where it becomes the only thing that matters. ## Who Sees Ads in ChatGPT (and Who Doesn't) Ads only reach the lower tiers. According to OpenAI's [Help Center documentation](https://help.openai.com/en/articles/20001047-ads-in-chatgpt), ads can appear for users on the Free and ChatGPT Go plans, whether or not they are logged in. As of the current rollout, everyone on a higher paid tier is exempt.
PlanSees ads?Monthly price
FreeYesFree
GoYes~$8
PlusNo~$20
ProNo~$200
Business / Enterprise / EducationNoVaries
A couple of limits apply even inside the ad-supported tiers. OpenAI [does not show ads to users it believes are under 18](https://www.stackadapt.com/resources/blog/how-to-advertise-on-chatgpt), and during the current test ads do not appear in Temporary Chats or right after you generate an image. If you are on Go and surprised to see ads, that is by design: Go is a cheaper subscription, not an ad-free one. You do not have to pay to escape them, though: the ad-free mode, which trades ads for fewer messages a day and no access to tools like image generation, is a Free-plan option, so a Go user would drop back to Free to use it. The other way out is to move up to Plus or higher. For the full plan breakdown, see our [ChatGPT pricing guide](https://geotoolbox.ai/blog/chatgpt-pricing). The European Union came last, and on its own terms. On August 15, 2026, OpenAI emailed Free and Go users across the European Economic Area and Switzerland that ads would start appearing there later in the month, non-personalized at first, with personalization requiring a separate opt-in. That phased, consent-first design is how the rollout works around GDPR consent rules and the Digital Services Act's ad-transparency requirements, which are the likely reason the EU waited longest. ## Where Ads Appear and What They Look Like Ads sit below the answer, never inside it. When ChatGPT finishes replying, a sponsored unit can appear beneath the response, clearly labeled and visually separated from the text the model wrote. It reads more like a native card than a banner. The creative is compact. Per StackAdapt, which runs ChatGPT ads as a partner, [each ad carries the brand's name, a favicon, a headline, a short description, an image, and a link to the landing page](https://www.stackadapt.com/resources/blog/how-to-advertise-on-chatgpt), built around a 256-by-256 pixel square image (StackAdapt cites headline and body caps around 30 and 60 characters, though other spec guides list lower numbers, so confirm the current limits in Ads Manager). A response usually shows a single sponsored unit, but OpenAI has been testing units that group more than one advertiser, and product-feed formats let retailers surface several items at once. Be careful with older write-ups here. A lot of early-2026 coverage described "sidebar" or "companion display" ad units and a menu of three or four formats. That was speculation from launch week. What OpenAI's [own advertiser docs](https://developers.openai.com/ads) actually describe are sponsored units below the answer plus product-feed variants, not a Bing-style sidebar. ## How ChatGPT Decides Which Ads to Show The core signal is your current conversation. OpenAI's system reads what you are asking about, runs a brand-safety check, finds ads whose topic and landing page match, and picks the most relevant one. Advertisers do not bid on exact-match keywords the way they do in Google Search. Instead they supply "context hints," plain-language descriptions of the topics and situations where their product fits, and OpenAI decides the final match. Geographic targeting is available too, at the country level and, in the US, down to state, metro area, and ZIP. Targeting is contextual by default, but it is not only contextual. If you have ad personalization turned on, OpenAI says it [may also use your past chats, memory, and previous ad interactions](https://help.openai.com/en/articles/20001047-ads-in-chatgpt) to make ads more relevant. That is behavioral targeting, and a good deal of the early-2026 coverage that called ChatGPT ads "purely contextual, no user profiles" simply got it wrong. What advertisers never receive is your raw material. Per OpenAI, they do not get your chats, history, memory, name, email, precise location, or IP address. The matching happens inside OpenAI, which acts as the middleman. You can turn personalization off, clear your ad data, or use a Temporary Chat to keep a session out of it. ## Do ChatGPT Ads Change the Answers? This is the question most users actually care about, and OpenAI's answer is no. Its documentation states that [ads do not influence the answers ChatGPT gives](https://help.openai.com/en/articles/20001047-ads-in-chatgpt), that the ad system is separate from the model that writes responses, and that advertisers cannot shape, rank, or alter what ChatGPT says. An ad appearing below an answer is not an endorsement of that brand by ChatGPT. Whether you take that on faith is up to you, and plenty of users are skeptical, especially given that OpenAI spent years calling in-product ads a last resort before reversing course. The incentive to blur the line grows as ad revenue does, so the separation is worth watching, not just trusting. But the architecture is what matters today, and it cuts in a useful direction. Because the answer and the ad are generated separately, ad spend cannot move you into the answer itself. The brands named in a response come from what the model learned and retrieved, and paying for an ad does not change that. We come back to what that means for visibility [below](#what-chatgpt-ads-mean-for-your-organic-visibility). For the mechanics of how ChatGPT chooses which sources to name, see [how ChatGPT cites sources](https://geotoolbox.ai/blog/chatgpt-citations). ## What ChatGPT Ads Cost (for Advertisers) There is no published rate card, and the pricing has moved fast. What has stayed constant is the buying model: you pick an objective, reach priced on cost per thousand impressions (CPM), clicks priced on cost per click (CPC), or a conversion-optimized version of CPC, set a maximum bid, and OpenAI runs a relevance-weighted second-price auction. What changed is the barrier to entry.
StageDateWhat it cost to get in
Launch pilotFeb 2026~$60 CPM, $200,000 minimum (enterprise only)
Minimum loweredApr 2026~$50,000 minimum
Self-serve opensMay 2026No minimum spend
NowAug 2026Auction: CPM, CPC, or conversion-optimized CPC, no minimum
The launch numbers, roughly $60 CPM behind a $200,000 commitment, are what made early headlines. By May, when [OpenAI opened self-serve buying](https://developers.openai.com/ads), the minimum was gone. Ad Age reported in July that the platform had [shifted from an impressions business toward a clicks business](https://adage.com/technology/ai/aa-chatgpt-ads-ad-tech-what-marketers-still-need/) as it matured. Reported CPCs since then land in the low single digits to low teens depending on the category, though those figures come from marketers, not OpenAI, so treat them as directional. Is it worth it? Early advertisers are mixed. The audience is large but skews toward users who did not pay to remove ads, which is a real consideration if you sell to buyers who would. The reporting is still thinner than Google or Meta, with no conversation-level visibility into which prompts triggered your ad. One more limit worth knowing before you plan a budget: OpenAI applies content and placement policies that keep ads out of sensitive contexts, so [regulated categories may be limited or unavailable](https://www.stackadapt.com/resources/blog/how-to-advertise-on-chatgpt) for now. For most brands this is a channel to test with a small budget, not to bet on yet. ## How to Advertise on ChatGPT If you want to run one, the entry point is [OpenAI's Ads Manager](https://developers.openai.com/ads) at ads.openai.com, or a technology partner if you already buy media through one. The structure will feel familiar: a campaign holds your budget and objective, an ad group holds your targeting and context hints, and the ad itself holds the headline, description, and image. The short version of a first campaign: create the campaign and pick reach or clicks, write context hints that describe the conversations where your product belongs, set geographic targeting, upload creative within the current character limits, and add the measurement pixel or Conversions API so you can see what happens after the click. Add UTM parameters to the landing page and the traffic will show up in your normal analytics. This is not a deep-optimization platform yet, so most of the early skill is in writing hints and creative that match how people actually phrase questions. ## What ChatGPT Ads Mean for Your Organic Visibility There is a strategic point buried under the ad mechanics. Everyone on the paid tiers, Plus, Pro, Business, Enterprise, and Education, browses ChatGPT ad-free. If your buyers skew toward those tiers, and for a lot of B2B, premium, and developer products they do, the people you most want to reach never see an ad, and the only way to appear for them is to be named inside the answer itself. If you sell broad consumer products, the free ad tier may be the bigger prize, so weigh both. The ad auction, at least, does not buy your way into the answer. Because the ad system and the model are separate under OpenAI's current design, the brands named in a response come from what the model learned and retrieved, not from ad campaigns. So ads and organic visibility are not substitutes. They are layered. One is a recommendation layer, where you earn a mention inside the response across every tier. The other is a consideration layer, where a sponsored card can appear below the answer for Free and Go users. Ads buy the second. The first you earn by being citable. The one paid-adjacent exception is OpenAI's separate commerce track, where merchant product feeds can surface items in shopping flows, but that is a different surface from the chat answer.
![A ChatGPT reply shown as stacked layers: the organic answer on top, labeled every tier sees this and naming brands the model chose, and a labeled Sponsored ad card below it, labeled only Free and Go see this, generated by a separate system.](/blog/chatgpt-ads/chatgpt-ads-organic-vs-paid-layers.png)
The answer and the ad are separate systems. Ads can appear below the response for Free and Go users, but ad spend does not change who the model names inside the answer.
That reframes the ChatGPT ad launch as less of a threat to organic strategy and more of a confirmation of it. If OpenAI is now charging brands to appear near answers, the answers themselves are clearly where the attention is. Getting recommended in those answers is a different discipline, [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo). We cover the practical steps in [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search), and if you are deciding how to divide effort between classic search and AI, [GEO vs SEO](https://geotoolbox.ai/blog/geo-vs-seo) walks through the split. At geotoolbox we build tools for exactly this layer, checking whether AI engines can crawl your site and whether they actually cite you when it counts. The ad platform does not change that work. It makes it matter more. **Further reading:** [SEO for ChatGPT: How to Get Cited in ChatGPT Search](https://geotoolbox.ai/blog/seo-for-chatgpt) is the companion to this piece. It covers exactly how to earn a place in the answer, which is the visibility ads cannot buy. ## Frequently Asked Questions ### Do I have to pay to get rid of ads in ChatGPT? Not necessarily. Ads appear on the Free and Go plans, and moving up to Plus, Pro, Business, Enterprise, or Education removes them. But you do not have to pay: the ad-free mode, which trades ads for fewer messages a day and no access to tools like image generation, is a Free-plan option (a Go user would move back to Free to use it). ### How do I turn off or block ChatGPT ads? You have a few levers. To keep ads but make them less targeted, turn off ad personalization in settings, which stops OpenAI from using your past chats and ad history, and clear your ad data; ads still appear, just less personalized. To skip ads in a single session, use a Temporary Chat, which currently shows none. To remove them across the board, switch to the ad-free mode or upgrade to a paid tier. ### Are my ChatGPT conversations shared with advertisers? No. OpenAI says advertisers do not receive your chats, chat history, memory, name, email, precise location, or IP address. The ad matching happens inside OpenAI, which selects and serves the ad without handing your data to the brand. ### Will ads bias or change ChatGPT's answers? OpenAI states that ads do not influence the answers and that advertisers cannot shape, rank, or alter what the model says. The ad system runs separately from the model that writes responses. An ad below an answer is a paid placement, not a recommendation from ChatGPT. ### How much do ChatGPT ads cost to run? There is no fixed rate. Ads are bought through an auction on a CPM, CPC, or conversion-optimized CPC basis, and OpenAI does not publish official prices. The pilot launched around a $60 CPM with a $200,000 minimum in February 2026, but the minimum was removed when self-serve buying opened in May. Reported click costs since then run from a few dollars into the low teens depending on the category. ### Which countries have ChatGPT ads? Ads started in the United States in February 2026 and, as of an August 11, 2026 update, are also live in Canada, the United Kingdom, Australia, New Zealand, Japan, South Korea, Mexico, and Brazil. The European Economic Area and Switzerland come next: OpenAI announced on August 15, 2026 that ads would reach Free and Go users there later that month, non-personalized at first and with a separate opt-in before any personalization. Availability keeps changing, so the [current list lives in OpenAI's documentation](https://help.openai.com/en/articles/20001047-ads-in-chatgpt). ### Can my brand get recommended in ChatGPT without paying for ads? Yes, and it is a different job from advertising. Being named inside ChatGPT's answer depends on whether the model can find and trust your content, not on ad spend. That is what generative engine optimization addresses, and it is the only way to reach the paid tiers that never see ads. ## The Bottom Line ChatGPT ads are real, they are labeled, they sit below the answer, and they reach Free and Go users on a contextual auction that is still maturing. If you sell something, they are worth a small test. But the answer above the ad is where the lasting value is, and the ad auction does not buy a place in it. If you want to know whether AI engines can even reach your site, and whether ChatGPT is citing you when people ask about what you do, run a free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) and start with what the model can actually see. ## Sources - Our Approach to Advertising and Expanding Access - OpenAI, January 16, 2026 - `openai.com/index/our-approach-to-advertising-and-expanding-access` - Testing Ads in ChatGPT - OpenAI, February 9, 2026 - `openai.com/index/testing-ads-in-chatgpt` - Ads in ChatGPT (Help Center) - OpenAI - `help.openai.com/en/articles/20001047-ads-in-chatgpt` - Ads (advertiser documentation) - OpenAI Developers - `developers.openai.com/ads` - ChatGPT Rolls Out Ads - TechCrunch, February 9, 2026 - `techcrunch.com/2026/02/09/chatgpt-rolls-out-ads` - How ChatGPT Advertising Works Now - Ad Age, July 22, 2026 - `adage.com/technology/ai/aa-chatgpt-ads-ad-tech-what-marketers-still-need` - How to Advertise on ChatGPT - StackAdapt, May 28, 2026 - `stackadapt.com/resources/blog/how-to-advertise-on-chatgpt` --- ## LLM SEO: What It Is, What's Different, and How to Do It (2026) > LLM SEO means getting your brand cited by ChatGPT, Claude, and Perplexity. What changes vs traditional SEO, what Google says to skip, and how to do it. - Canonical: https://geotoolbox.ai/blog/llm-seo - Published: 2026-08-15 · Updated: 2026-08-15 LLM SEO is one of the most oversold terms in marketing right now, and one of the most misunderstood. Strip out the hype and a straightforward discipline sits underneath: getting your brand named when ChatGPT, Claude, Gemini, and Perplexity answer a question. This guide separates what works from what vendors are selling, including what Google itself says you can safely ignore. ## What LLM SEO Is (and the Confusion to Clear First) LLM SEO is the practice of structuring your content and brand presence so large language models like ChatGPT, Claude, Gemini, and Perplexity mention and cite you when they answer a question. The goal is not a blue-link ranking. It is being the source the model pulls from. Before anything else, clear up the phrase, because it is ambiguous in a way that sends people down the wrong path. Two people search "LLM SEO" and mean opposite things: - **Optimizing to appear in LLMs.** You want ChatGPT to name your brand when someone asks it for a recommendation. That is what this article is about. - **Using an LLM to do your SEO.** You want ChatGPT to write meta descriptions or draft content faster. That is a writing-workflow question, not a visibility one. So, can ChatGPT do SEO? It can help you produce and audit content, but "LLM SEO" as a discipline means the first thing: earning a place inside the answer. Keep the two apart and the rest of the field stops sounding like noise. That noise includes a pile of acronyms. They all describe the same underlying job, getting your brand surfaced and cited inside AI-generated answers, from different angles.
TermWhat it emphasizes
LLM SEOThe models specifically (ChatGPT, Claude, Gemini)
LLMO (LLM optimization)Same as LLM SEO, phrased around "optimization"
GEO (generative engine optimization)Generative answer engines broadly
AEO (answer engine optimization)Direct-answer surfaces like AI Overviews and featured snippets
AI search optimizationThe plain-English umbrella
Vendors will sell you separate budgets for each. Do not buy it. We map the full vocabulary in [what LLMO is](https://geotoolbox.ai/blog/what-is-llmo) and the [GEO vs AEO vs SEO breakdown](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo); pick one word for your deck and move on. What matters is the work, and it starts with what changed under the hood. ## What Changes When a Language Model Ranks You: Training vs Retrieval Here is the one idea most LLM SEO guides blur, and getting it straight resolves half the arguments in the field. A language model can surface your brand through two completely separate paths, and they run on different clocks.
![LLM SEO's two paths: training, which moves in model generations, and retrieval, which moves in days.](/blog/llm-seo/llm-training-vs-retrieval-paths.png)
The two ways a language model can surface your brand, and why they move at completely different speeds.
**The training path.** During pretraining, a model reads a large slice of the public web and absorbs patterns about who is associated with what. If enough sources describe you as a leading option in your category, that association becomes part of the model's baseline knowledge. You cannot edit this directly, and it moves at the speed of model releases: months, sometimes longer. This is why a brand that was quiet a year ago can still be missing from a model's "from memory" answer today. **The retrieval path.** When a model runs a live search to answer you, a separate system fetches current pages, and the model writes its answer from what it just pulled. This is [retrieval-augmented generation](https://geotoolbox.ai/blog/what-is-rag), and it moves in hours to days, depending on how often the retrieval crawler visits you. Publish a strong, reachable page and it can be cited in a live answer long before any model is retrained. Our breakdown of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) walks through the full loop. Once you see the two paths, the common confusions dissolve. "I got cited in ChatGPT within a day" and "AI visibility takes months to build" are both true: the first is retrieval, the second is training. Freshness matters most on the retrieval path, especially for time-sensitive queries, and barely at all on the training path. Two more shifts follow from this. There is no ranked list of ten blue links to climb; there is an answer, and you are either named in it or you are not. And the win is a citation, not a click, which changes how you prove the work is paying off (more on that below). ## LLM SEO vs Traditional SEO (and GEO, AEO) Start with the uncomfortable part for anyone selling LLM SEO as a brand-new discipline: most of it is traditional SEO. If a crawler cannot fetch and render your page, no retrieval layer can pull it into an answer. Crawlability, fast rendering, server-side HTML, and clean information architecture are the entry ticket, not a bonus. Google says this plainly. Its own guidance states that "the best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems," per [Google Search Central's AI-features guidance](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide). For Google's AI Overviews and AI Mode specifically, LLM SEO is SEO. What changes is the last stretch: how the answer is assembled and what gets you named in it.
DimensionTraditional SEOLLM SEO
What you optimizeA page, for a keywordA page and your off-site footprint, for a question
The winA ranked blue linkA citation or mention inside the answer
Who decidesOne ranking system (Google)Several engines with different indexes and rules
Authority signalBacklinksBacklinks plus unlinked mentions and third-party consensus
How you measureRankings, clicks, impressionsMention rate, citation rate, share of voice
TimelineWeeks to monthsDays (retrieval) to model generations (training)
Most of the difference sits in two rows. First, different engines do not reprint Google's top ten. They use their own indexes and source-selection rules, so ranking first on Google helps but does not guarantee you show up in ChatGPT (which draws on a mix of Bing and its own index) or Perplexity (which runs its own crawler). Second, authority stops being only about links to your site. A model builds its picture of you from what the rest of the web says, which is why the work leaks off your own domain (more on that in the playbook). If you are deciding where to spend, this is a rebalancing of one budget, not a second one. We lay out [how to split your GEO budget](https://geotoolbox.ai/blog/geo-vs-seo), and cover the definitions in [what generative engine optimization is](https://geotoolbox.ai/blog/what-is-geo) and [what answer engine optimization is](https://geotoolbox.ai/blog/what-is-answer-engine-optimization). The tactics below are where the new part lives. ## The LLM SEO Playbook: Five Levers That Move Citations Five levers do most of the work. None of them is a trick. The order matters, because the first one gates all the others.
LeverWhy it works with an LLMGo deeper
1. ReachabilityIf AI crawlers can't fetch the page, the retrieval path never sees itAI crawlers
2. Answer-first structureA direct answer under a clear heading is a clean passage to lift into a responseStructuring for extraction
3. Citable substanceStatistics, citations, and quotes give a model concrete, attributable materialE-E-A-T for AI
4. Off-site presenceModels build your entity picture from what other sites say about youEntity SEO
5. FreshnessRetrieval leans on current pages for time-sensitive questions; evergreen authority ages more slowlyAI content optimization
**Citable substance is the one with a controlled study behind it.** A widely cited [academic GEO study](https://arxiv.org/abs/2311.09735) tested content changes against a benchmark of real search queries and found that adding citations, quotations from credible sources, and statistics were the top-performing methods, with visibility gains of roughly 30 to 40 percent on its main metric. Plain-language, readable phrasing helped too. Keyword stuffing, the old SEO reflex, did not. So the highest-impact edit is often the least glamorous: put a real number, a named source, or a direct quote next to each claim. **Answer-first structure needs a caveat**, because Google is explicit that you do not need to break content "into tiny pieces for AI to better understand it." That is true for Google's own AI features, which run on its core ranking systems. But for standalone retrieval engines like Perplexity and ChatGPT search, a clear answer immediately under a descriptive heading is simply a cleaner passage to extract. Write that way because it serves the reader, not because you are feeding a machine. The moment the structure hurts a human, you have gone too far. **Off-site presence is the lever that breaks SEO habits.** In classic SEO, the unit of work is your page. In LLM SEO, a large share of the work is how you are represented on other people's pages: review directories, "best tools" listicles, comparison articles, YouTube videos and their transcripts, and community threads on Reddit and Quora. Models build a picture of your reputation from that chorus, and unlinked mentions can matter, not just links. This is closer to digital PR than to on-page optimization, and it is the part you control least. Reddit deserves a specific warning, because "just post on Reddit" is the most-repeated tactic in this space and the fastest way to get burned. Communities do corroborate a brand, but manufacturing consensus with a throwaway account is easy to spot and gets the account banned. Participate honestly or not at all. And because AI search often [fans one question out into many](https://geotoolbox.ai/blog/query-fan-out), the goal off-site is to be present across the sub-questions a buyer actually asks, not to plant one link in one thread. For the step-by-step version of all five, including exactly what to change on a page, follow our [playbook for optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search). The rest of this guide handles the parts that playbook assumes you have already sorted: whether the engines can reach you at all, whether any of this is working, and how you would know. ## Can AI Even Reach You? Reachability You Can Verify This is the failure that wastes the most effort, because it is silent. You can nail every tactic above and still be invisible if the crawler that feeds the answer never reaches your page. Most guides list "check your robots.txt" and move on. The trap is more specific than that. Different bots do different jobs, and blocking the wrong one has very different consequences. The major AI operators each run several, and the jobs differ (OpenAI documents its own set in its [crawler documentation](https://developers.openai.com/api/docs/bots)):
User-agentOperatorJobBlock it and...
GPTBotOpenAIGathers training dataYou opt out of training data, not live answers
OAI-SearchBotOpenAIIndexes pages for ChatGPT searchYour pages stop being cited in ChatGPT search
Google-ExtendedGoogleGemini training controlYou opt out of Gemini training (not Google Search)
PerplexityBotPerplexityIndexes pages for Perplexity answersYour pages stop being cited in Perplexity
Claude-SearchBotAnthropicIndexes pages for Claude searchYour pages stop being cited in Claude
One caveat that trips people up: the user-action fetchers (ChatGPT-User, Perplexity-User, Claude-User) grab a page when a person asks for it, and they do not all follow robots.txt. OpenAI says its rules may not apply to ChatGPT-User, and [Perplexity's user fetcher generally ignores robots.txt](https://docs.perplexity.ai/docs/resources/perplexity-crawlers), while [Anthropic says Claude-User honors it](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler). So a robots.txt block is only fully reliable for the automated crawlers above. The common and costly mistake is blocking the retrieval crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot) while thinking you only opted out of training. Your robots.txt "looks fine," and you have quietly stopped being cited. The gate you did not set is worse. In July 2025, [Cloudflare began blocking AI crawlers by default for new domains](https://www.cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large/), calling itself the first major infrastructure provider to do so. Its rules have kept shifting since: under the [2026 category controls](https://blog.cloudflare.com/content-independence-day-ai-options/), from September 15, 2026 new domains block Training and Agent crawlers by default on ad-displaying pages while leaving Search crawlers allowed. The takeaway is not the exact default of the month; it is that if your site sits behind a CDN or WAF, an edge policy you never chose can decide which AI bots reach you. Cloudflare's own [June 2024 data](https://blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click/) found AI bots reached about 39 percent of the top million properties on its network, while just 2.98 percent of the top million had any AI-bot controls in place at all. So verify, do not assume. Read your robots.txt line by line against the user-agents above, and our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) does exactly that: it shows which of 34 AI crawlers your robots.txt allows or blocks, with the exact line to fix. That covers the robots.txt layer only, though. To catch the WAF and CDN gates, grep your server logs for the fetching crawlers (GPTBot, OAI-SearchBot, PerplexityBot, Claude-SearchBot) to confirm they are getting a 200, and diff your raw HTML against the JavaScript-rendered version, since a crawler that does not run your JavaScript sees only the raw response. Across the sites we audit for AI visibility, a stale robots.txt rule or an edge default is a common and easily-missed reason a reachable-looking page never gets cited, and it is the cheapest thing to rule out first. This reachability-first logic is also why we treat [llms.txt](https://geotoolbox.ai/blog/llms-txt) as a distraction for AI visibility. No major AI engine has confirmed using it as a citation signal, and an [Ahrefs study of 137,210 domains](https://ahrefs.com/blog/llmstxt-study/) found that 97 percent of the llms.txt files it saw received zero requests in May 2026. The file does have a use in some developer tooling, where coding agents can pull it to fetch docs, but that is not the same job as getting cited in an answer. ## Does LLM SEO Actually Work? An Honest Look at the Evidence Yes, with a few caveats. The market is not in doubt: Alphabet's Sundar Pichai said Google's AI Overviews passed [2.5 billion monthly users](https://blog.google/innovation-and-ai/sundar-pichai-io-2026/) at Google I/O in May 2026, and AI Mode passed 1 billion. The question is not whether AI answers matter. It is whether your effort moves your slice of them. The strongest number in the field is the GEO study's "up to 40%" visibility lift, which we cited above. Read what it actually measured: a gain on a research-built pipeline over the top Google results in 2023 and 2024, scored on a benchmark visibility metric, not commercial traffic to a business running on production ChatGPT or Perplexity. It tells you which content changes a model responds to. It does not promise you 40 percent more customers. Anyone quoting it as a business outcome is stretching it. The most-repeated actionable claim in this space is that AI citations drop off sharply once content is more than about three months old. It gets restated across guide after guide, and we could not find a single named dataset behind it. It is also, conveniently, repeated most often by companies selling the tracking and refresh services it justifies. Treat it as a plausible hypothesis about the retrieval path, not a measured law. The pattern to plan for in the near term is that your visibility goes up while your click traffic stays flat or dips. Being named in an answer is a branding and consideration win, and it often does not send a click at all. If you report [AI SEO](https://geotoolbox.ai/blog/ai-seo) to a client or a boss using only sessions and clicks, a genuine win can look like a failure. The fix is to set the terms before you start: report citation share and branded-search lift alongside sessions, and baseline all three at the outset so the trade is visible rather than hidden. So when is LLM SEO not worth it? If you have not fixed basic crawlability and indexing, or if you cannot commit to the off-site and content work over months, do the SEO fundamentals first and revisit. The harder question is whether your buyers use AI tools to research your category yet, and you can check it rather than guess: look at the share of traffic in GA4's AI Assistant channel (a lower bound, since app visits hide in Direct), spot-check a handful of buying-intent prompts in ChatGPT and Perplexity to see who gets named, and ask a few customers how they found you. LLM SEO amplifies a solid foundation. It does not substitute for one. For where the wider market sits, our [state of AI search report](https://geotoolbox.ai/blog/state-of-ai-search-2026) lays out the data. ## How to Measure LLM SEO (Without Fooling Yourself) Measurement is where LLM SEO gets hard. The core problem is that these systems are non-deterministic: ask the same model the same question twice and you can get two different answers, with different sources cited. A single check tells you almost nothing, which is why our [guide to measuring AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) is built around repetition rather than one-off lookups. So build a method instead of spot-checking: 1. Write a fixed, representative set of prompts your buyers ask, and freeze it. Changing the prompts changes the results, so the set has to stay stable to be a baseline. 2. Run each prompt several times per check, across each engine you care about, and record how often you are mentioned and how often you are cited with a link. One run is noise; the rate across runs is the signal. 3. Log which URL got cited, not just that you appeared. The page that gets pulled tells you what is working. 4. Hold a fixed competitor set and track your share of voice against it over time, so you are measuring movement, not a single snapshot. 5. Report per engine. ChatGPT, Gemini, and Perplexity draw from different indexes, so a blended number hides where you are winning and losing. Keep three things separate while you do this. A **mention** is the model naming you in prose. A **citation** is a linked source. A **referral** is a click that actually lands on your site. They are not the same, and conflating them is how reports mislead. That last one has a trap. The native ChatGPT, Claude, and Perplexity apps often do not pass a referrer, and GA4 files any referrer-less visit as **Direct**, so app-driven AI visits quietly land in Direct rather than showing up as AI referrals. GA4 added a built-in AI Assistant channel in 2026 that groups recognized engines automatically, which helps, but it is still referrer-based, so any AI visit without a referrer or a tag still falls to Direct. The one durable signal is a UTM in the link itself: ChatGPT appends `utm_source=chatgpt.com` to its citations, and a query-string tag survives referrer stripping. Treat any referral-based number as a floor, and pair it with prompt-based tracking that does not depend on a click happening at all. Our guides on [running an AI visibility audit](https://geotoolbox.ai/blog/ai-visibility-audit) and [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) go deeper. One honest note, since we build a tool in this category. Every AI visibility tracker, ours included, samples a prompt set and estimates; none of us can see the real query volumes inside ChatGPT, and results shift run to run. A tracker's job is to make that sampling disciplined and repeatable, not to hand you a precise number that does not exist. Anyone promising exact AI query counts is selling certainty the systems do not expose. ## Common LLM SEO Mistakes Most wasted effort in LLM SEO comes from a short list of avoidable errors: - **Blocking the retrieval crawler by accident.** The quiet failure from earlier: you mean to opt out of training, but block a retrieval crawler too and drop out of live answers without noticing. Verify per user-agent. - **Mass-producing AI-generated content.** The web is flooding with synthetic pages, which makes original data, first-hand testing, and named expertise more valuable, not less. Publishing more of the same slop moves you the wrong way. - **Confusing AI-only markup with normal structured data.** Google is explicit that you do not need AI text files (its wording for files like llms.txt), AI-only markup, or extra schema.org markup to appear in its AI features. Standard structured data is a different thing: Google still recommends keeping it for rich results, and it stays ordinary SEO hygiene. Competitors pointing out that "nearly every cited page has schema" are describing a correlation, not proof that schema earns the citation. Keep your normal schema, skip the AI-specific inventions, and see [what schema does and doesn't do for AI search](https://geotoolbox.ai/blog/schema-markup-for-ai) for the detail. - **Gating your best content behind a login and expecting citations.** If a crawler cannot read it, a model cannot cite it. The middle path is to expose a substantive, self-contained summary that can stand on its own and gate the rest, or use Google's documented markup for paywalled content. What you cannot do is serve AI bots the full text while blocking human readers; that is cloaking, and it is against search guidelines. - **Chasing every engine at once.** Pick the two or three your buyers actually use and go deep. Spreading thin across eight surfaces wins none of them. - **Naming yourself into a collision.** Inventing a branded category ("we're a marketing intelligence engine") can collide with an established term in the training data and confuse retrieval about what you even are. Describe yourself in words the model already associates with your category. ## Frequently Asked Questions ### Is SEO dead now that AI answers the question directly? No. AI answers lean heavily on the same search indexes and live web, so crawlability, indexability, and quality content are still the foundation. What is dying is the assumption that a ranking equals traffic; a chunk of that attention now resolves inside the answer without a click. The work is shifting rather than ending. ### Can you get cited by AI without ranking well in Google? Yes, and it happens often. Answer engines use their own indexes and can pull a well-structured, reachable page into an answer even if it sits on page two of Google. Strong Google rankings help your odds, but they are neither required nor a guarantee. ### Does ChatGPT cite the same pages that rank number one in Google? Sometimes, but do not count on it. The overlap is partial and swings with the query: informational questions tend to pull more from the familiar authorities, while commercial and comparison prompts often surface listicles, forums, and review pages that never ranked first. Ranking well is a positive signal, but plan for the AI surface as its own target rather than assuming your Google winners carry over. ### Do I need an llms.txt file for LLM SEO? No. No major AI engine has confirmed using llms.txt as a citation signal, a large Ahrefs study found 97 percent of published files got zero requests, and Google says directly that you do not need AI text files to appear in its AI features. It has a real use in developer tooling, but for getting cited your time is better spent confirming AI crawlers can reach your pages. ### How long does LLM SEO take to show results? It depends on the path. On the retrieval path, a reachable, well-structured page can be cited in live answers within days. On the training path, becoming part of a model's baseline knowledge takes model generations, often many months, and depends on broad third-party mentions you cannot rush. ### Can I stop AI from recommending my competitor instead of me? Not directly. You cannot edit a model, and you cannot remove a competitor from its answers. What you can do is strengthen the signals the model reads about you, which means earning more third-party mentions, reviews, and citable content so the balance of evidence shifts in your favor over time. ### What do I do if an AI describes my business incorrectly? Fix the sources it is reading, since you cannot edit the model. Correct outdated facts on your own pages, then work on the third-party pages, directories, and profiles the model pulls from, because a confident wrong answer often traces back to stale or conflicting information out on the web, though it can also be a plain model fabrication. Our guides on [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations) and [tracking AI brand sentiment](https://geotoolbox.ai/blog/ai-brand-sentiment) cover how to find and unwind these. ## Where to Start Underneath the acronyms, LLM SEO is mostly disciplined SEO, plus digital PR, plus one thing the old playbook never had to worry about: whether the machines assembling the answer can reach you at all. The teams that win are not the ones buying a separate "GEO budget." They are the ones doing the fundamentals well and then covering the new ground, reachability and off-site presence, with the same rigor. Start with the gate, because it is the cheapest thing to check and a common reason good work goes nowhere. Run your site through our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) to see which AI bots your robots.txt allows, then check your server logs and CDN for the ones that slip past it. If a retrieval crawler is blocked, fix that first. Everything else in this guide only pays off once the answer engines can see you. ## Sources - GEO: Generative Engine Optimization - Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan, Deshpande, 2023-2024 - `arxiv.org/abs/2311.09735` - AI features and your website (best practices for SEO) - Google Search Central, updated 2026-07-10 - `developers.google.com/search/docs/fundamentals/ai-optimization-guide` - Overview of OpenAI crawlers (GPTBot, OAI-SearchBot, ChatGPT-User) - OpenAI - `developers.openai.com/api/docs/bots` - Perplexity crawlers (PerplexityBot, Perplexity-User) - Perplexity - `docs.perplexity.ai/docs/resources/perplexity-crawlers` - Does Anthropic crawl the web, and how to block it (ClaudeBot, Claude-User, Claude-SearchBot) - Anthropic - `support.claude.com/en/articles/8896518` - Cloudflare just changed how AI crawlers scrape the internet (blocking AI crawlers by default) - Cloudflare, 2025-07-01 - `cloudflare.com/press/press-releases/2025/cloudflare-just-changed-how-ai-crawlers-scrape-the-internet-at-large` - Your site, your rules: new AI traffic options (2026 category controls) - Cloudflare, 2026-07-01 - `blog.cloudflare.com/content-independence-day-ai-options` - Declaring your AIndependence: block AI bots with one click - Cloudflare, 2024-07-03 - `blog.cloudflare.com/declaring-your-aindependence-block-ai-bots-scrapers-and-crawlers-with-a-single-click` - Do LLMs.txt files actually work? (137,210-domain study, May 2026 data) - Ahrefs, 2026-06-15 - `ahrefs.com/blog/llmstxt-study` - Sundar Pichai remarks, Google I/O (AI Overviews 2.5B monthly users, AI Mode 1B) - Google, 2026-05 - `blog.google/innovation-and-ai/sundar-pichai-io-2026` --- ## Semantic SEO: How to Write for Meaning, Not Keyword Strings > Semantic SEO means optimizing for topics, entities, and meaning instead of exact-match keywords. Here's the retrieval mechanism behind it and how to do it. - Canonical: https://geotoolbox.ai/blog/semantic-seo - Published: 2026-08-15 · Updated: 2026-08-15 Semantic SEO is the practice of optimizing content for topics, entities, and meaning rather than exact-match keyword strings. It is not a new trick bolted onto search. It is what search quietly became, and what AI answer engines now lean on heavily. The confusing part is that most explanations stop at "optimize for topics, not keywords" and leave the mechanism a black box. This one opens the box: how engines turn your writing into meaning, why they cite one passage and ignore the rest of the page, and what that means for the choices you make while writing. ## What Is Semantic SEO? **Semantic SEO is optimizing content so search and AI engines understand its meaning, not just match its words.** You write for a topic and the entities inside it, cover the intent behind the query, and make the relationships between concepts clear. The engine's job is to resolve what your page is about; your job is to make that resolution easy. The shorthand is "things, not strings." A keyword is a string of characters. A concept or an [entity](https://geotoolbox.ai/glossary/entity-seo) is a thing that stays the same across every phrasing of it. Optimize for the string and you win one query. Optimize for the thing and you can surface for the many ways people ask about it. A fair question SEOs keep asking: is semantic SEO just good writing rebranded? Mostly, yes. The difference is that you now know the mechanism that rewards good writing, which tells you which advice has a reason behind it and which is superstition. Keywords are not dead either. They are still an input, a signal the engine reads. They are just no longer the finish line.
DimensionKeyword SEOSemantic SEO
Unit of optimizationA string of wordsA topic and its entities
Main signalsExact-match keywords, densityMeaning, intent, relationships, coverage
What you winOne ranked queryA cluster of related queries and AI citations
How you measureSingle-keyword rankTopic-level impressions and citations
Google's own documentation is explicit about this. Its [How Search Works page](https://www.google.com/search/howsearchworks/how-search-works/ranking-results/) describes a "sophisticated synonym system that allows us to find relevant documents even if they don't contain the exact words you used," and adds a line that should end the keyword-density debate on its own: when you search for "dogs," you "likely don't want a page with the word dogs on it hundreds of times." Meaning is the target. Repetition is not. ## Why Semantic SEO Matters More Now Classic search went semantic more than a decade ago. The milestones trace one direction: Hummingbird in 2013 reworked the engine around meaning, RankBrain added machine learning to interpret unfamiliar queries in 2015, BERT brought language understanding to Search in 2019 (the research paper landed in 2018), and MUM arrived in 2021 for specific search tasks. Each step moved the engine further from matching words and closer to understanding them. What changed recently is the stakes. AI answer engines (ChatGPT Search, Perplexity, Google's AI Overviews and AI Mode) do not show a page of blue links you can skim. They read the meaning of a question, pull the passages that answer it, and write an answer with a handful of citations. If your page is not understood at the level of meaning, it is much less likely to be pulled in. Semantic SEO stopped being an edge and became the baseline. If you are deciding how much of this to prioritize, our guide on the [SEO and AI budget split](https://geotoolbox.ai/blog/geo-vs-seo) puts numbers around it. It is also why the search volume for "semantic seo" looks flat while the practice matters more than ever. The demand did not shrink. It moved up a layer, into [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), where the same mechanics decide who gets cited. ## How AI Engines Actually Read Your Content This is the part that makes everything else make sense. An engine cannot compare passages the way you do. So it converts them into numbers. A [vector embedding](https://geotoolbox.ai/blog/vector-embeddings) is a list of numbers that captures what a passage means. Two passages about the same thing get vectors that sit close together; two unrelated passages end up far apart. Closeness is often measured with cosine similarity, and the useful property is that it works without shared words. A page that never says "cheap flights" can still be the closest match to "affordable airfare," because meaning, not vocabulary, decides the distance. Retrieval is where this gets used, and it tends to be hybrid rather than purely semantic. Production engines do not publish their exact stacks, but the well-documented pattern combines lexical matching, the classic keyword-and-term approach whose canonical treatment is Robertson and Zaragoza's [BM25 review](https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf), with dense matching on the embeddings above. This is why exact strings still matter for product names, error codes, and SKUs, and why "just write naturally" is incomplete advice. You need both halves.
![A query is embedded, matched to candidate passages by lexical and dense similarity, reranked, and one self-contained passage is cited in the answer.](/blog/semantic-seo/how-ai-engines-retrieve.png)
How an AI engine goes from your query to a cited passage: embed, retrieve on meaning and keywords, rerank, and build the answer from the passage that stands on its own.
When an AI engine answers, it does not read the whole web live. It retrieves a shortlist of passages, reranks them (a second, more careful scoring pass that reorders the shortlist), and writes a response grounded in the winners. This is [retrieval-augmented generation](https://geotoolbox.ai/blog/what-is-rag), the pattern named in a [2020 paper by Lewis and colleagues](https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html) that pairs a neural retriever with a text generator. Google AI Mode adds a twist called [query fan-out](https://geotoolbox.ai/blog/query-fan-out), where one question is expanded into several sub-questions and each runs its own retrieval. The consequence: **the engine builds its answer from the passage, even though the link points to the whole page.** It draws on the specific chunk that answered the sub-question and attributes it to your URL. So the unit that has to stand on its own is not your article, it is each passage inside it. A section that only makes sense after reading the ones above it is hard to lift out, and a passage that cannot stand alone is easy for an engine to skip. This is why [content chunking](https://geotoolbox.ai/blog/content-chunking) matters more than word count. For the full pipeline end to end, see [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work). ## Is LSI a Real Ranking Factor? No. "LSI keywords" is one of the most durable myths in SEO, and it is worth clearing up because it sends people in exactly the wrong direction. Google's John Mueller [said it flatly in 2019](https://www.seroundtable.com/google-lsi-keywords-27970.html): there is "no such thing as LSI keywords," and anyone telling you otherwise is mistaken. Google does not maintain a hidden list of synonyms you must sprinkle in. The confusion comes from a real technique with a similar name. Latent Semantic Indexing (or Analysis) is a genuine information-retrieval method from a 1990 paper by Deerwester, Dumais, Furnas, Landauer, and Harshman, published in the Journal of the American Society for Information Science. It uses singular value decomposition on a word-document matrix to find latent structure in a small, fixed corpus. It predates modern search and does not scale to the live web, and it is not what Google runs. Modern embeddings grew out of the same distributional intuition, but they use neural methods and web-scale infrastructure that classic LSA never had.
The "LSI keywords" mythWhat actually helps
A secret list of synonyms Google rewardsCovering the concepts and questions a topic genuinely implies
Sprinkle related terms to hit a densityAnswer the sub-questions a reader (and the engine) will have
More matching words means more relevanceClearer meaning and named entities mean more relevance
The useful idea buried under the myth is real: when you cover a topic properly, the related terms show up on their own, because you cannot explain a subject without them. That is a byproduct of depth rather than a checklist to pad. ## What You Control On-Page vs What You Earn Off-Site Semantic SEO splits cleanly into on-page and off-page work, and confusing them is where a lot of effort gets wasted. On your page, you control clarity: how well each passage stands alone, how you structure sections, how you link them, and what you declare about your [entities](https://geotoolbox.ai/blog/entity-seo) in [schema markup](https://geotoolbox.ai/blog/schema-markup-for-ai). Off your page, you earn recognition: an entity grows stronger in the [knowledge graph](https://geotoolbox.ai/glossary/knowledge-graph) when other sources describe you the same way, consistently. That consistency matters most when your brand name is also a common word, where an engine has to tell your entity apart from the dictionary meaning, and a clean, uniform external footprint is what lets it tell them apart. This distinction matters because schema gets oversold. Structured data is a claim you make about yourself; it is not proof, and it is not a ranking lever on its own. Google's own guidance on [AI features](https://developers.google.com/search/docs/appearance/ai-features) is blunt: "There are no additional requirements to appear in AI Overviews or AI Mode," and "no special schema.org structured data that you need to add." Schema helps engines parse and disambiguate what you already demonstrate. It does not manufacture authority you have not earned. So the mental model is simple: write and structure for the machine on-page, and build corroboration off-page. ## How to Do Semantic SEO Here is the workflow, step by step. ### Map the topic, not a keyword list Start from the job the reader is doing and the entities involved, not a spreadsheet of phrases. Find the sub-questions the way an engine will: mine Google's People Also Ask and related searches, expand each with the obvious what, why, and how follow-ups, read the Reddit and forum threads where people ask it in their own words, and note the entities your competitors mention that you do not. Those sub-questions, not keyword variants, are your outline. ### Cover the concept and answer the fan-out Write to satisfy the intent completely. If the engine fans one query into several, your page should answer the ones it is genuinely the right home for, each in its own place. This is how one page ends up surfacing for a cluster of related searches instead of a single term. ### Structure for passage retrieval Write in self-contained chunks. Put the answer first in each section, keep one idea per section, and make sure a passage makes sense if it is lifted out on its own. A section that opens with "it depends on what we covered above" fails, because it needs the paragraph above it to mean anything; the same point rewritten as "content loads faster when you cut render-blocking scripts" survives on its own and can be quoted. Question-style headings and a real FAQ help here, because they line up with the sub-questions engines retrieve against. ### Build the semantic graph with internal links Connect related pages with descriptive, contextual links so the engine can see the shape of your coverage. A pillar page supported by cluster pages, wired together deliberately, is how [topical authority](https://geotoolbox.ai/blog/topical-authority) is actually built, and it does not come from hitting a fixed post count. ### Declare entities, then earn them Add schema that describes your entities and their relationships (`about`, `mentions`, `sameAs`) to make them easier to parse, not as a ranking lever, then do the off-page work of getting described consistently elsewhere. For the broader playbook, our guide on [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) goes deeper on each step. ## How to Measure Semantic SEO Stop grading semantic work by single-keyword rank. The right signal is topic-level: in Google Search Console, look at how many distinct queries one URL brings impressions for, and whether cluster impressions are rising. One page ranking for many related questions is a good sign your coverage is working. Then measure the outcome that classic tools miss. Ranking well and being cited by an AI engine are separate results, and a page can do one without the other. Across the sites we scan for AI visibility at geotoolbox, the pages that get quoted are rarely the ones stuffed with keywords; they are the ones where a single passage answers a question completely. To see whether engines are actually pulling you into answers, you have to [track AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) directly and watch your [share of AI answers](https://geotoolbox.ai/blog/ai-share-of-voice) over time. Visibility is not the whole story, though, and the vanity-citation worry is a fair one. A citation is not a click, so track the outcome too: watch for AI-referral sessions in your analytics, and treat branded-query clicks in Search Console as a weak, confounded signal of whether being quoted is pulling people toward you. Some citations send traffic and some only build familiarity. You want to know which is which before you call the work a win. ## Frequently Asked Questions ### What is a semantic SEO example? A recipe page that covers not just "banana bread recipe" but why bananas need to be overripe, what to substitute for baking soda, how to tell when it is done, and how to store it. It targets the whole topic and its sub-questions, so it can surface for dozens of related searches and be quoted in an AI answer, rather than chasing one exact phrase. ### Is semantic SEO different from traditional SEO? It is an evolution of it, not a replacement. Traditional signals like content quality, links, and technical health still apply. Semantic SEO changes the unit of optimization from the keyword to the topic and its entities, and adds structure that makes individual passages retrievable by AI engines. ### Is Google a semantic search engine? Partly. Google understands meaning through systems like RankBrain and BERT, but retrieval is hybrid: it still uses lexical, keyword-based matching alongside [semantic search](https://geotoolbox.ai/glossary/semantic-search). That is why exact terms like product names and error codes still need to appear on the page, even though meaning drives most of the work. ### Do I need special schema markup for AI search? No. Google states there are no additional requirements and no special structured data needed to appear in AI Overviews or AI Mode. Schema helps engines parse and disambiguate your content, but it is not a separate ranking lever and it does not create authority you have not earned elsewhere. ### Does semantic SEO help me get cited in ChatGPT and Perplexity? Yes. Those engines retrieve passages by meaning, so covering a topic thoroughly and writing self-contained, extractable passages makes your content easier to pull into an answer. No writing pattern guarantees a citation, but this is what puts you in contention. ## The Takeaway There is no mystery to semantic SEO. Engines read for meaning, retrieve by similarity, and build their answers from the passage that stands on its own. Get the mechanism right and the tactics follow: cover the topic, structure for retrieval, declare your entities, earn corroboration, and measure at the topic and citation level. The one thing keyword tools cannot tell you is whether AI engines actually understand and cite your pages. That is the gap geotoolbox is built to close. Our [content analyzer](https://geotoolbox.ai/features/content-analyzer) scores how citable each passage is and how readable your page is to AI engines, and the free [AI Readiness checker](https://geotoolbox.ai/tools/ai-readiness) shows whether engines can reach and parse your content in the first place. Start with those checks. ## Sources - Google - How Search Works: Ranking results - `google.com/search/howsearchworks/how-search-works/ranking-results` - Google Search Central - AI features and your website - `developers.google.com/search/docs/appearance/ai-features` - Lewis et al. (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks - NeurIPS - `proceedings.neurips.cc/paper/2020` - Robertson & Zaragoza (2009), The Probabilistic Relevance Framework: BM25 and Beyond - `staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf` - Google's John Mueller on LSI keywords (2019) - Search Engine Roundtable - `seroundtable.com/google-lsi-keywords-27970.html` - Deerwester, Dumais, Furnas, Landauer & Harshman (1990), Indexing by Latent Semantic Analysis - Journal of the American Society for Information Science --- ## GEO vs SEO: How to Split Your Budget Between Search and AI > GEO vs SEO (and SEO vs GEO): what each optimizes for, when traditional SEO still wins, how to split the budget between them, and when to migrate to GEO. - Canonical: https://geotoolbox.ai/blog/geo-vs-seo - Published: 2026-08-14 · Updated: 2026-08-16 The real question behind GEO vs SEO is not which one to pick. It is how to divide a fixed budget now that AI answers sit between your pages and your audience. The two are not rivals; they are two surfaces built on one foundation, and the money you already spend on search is what makes AI citations possible in the first place. Most of this piece is about where the incremental dollar goes. One thing to clear up first, since the acronym is overloaded: GEO here means generative engine optimization, not geo-targeting or local search. ## The Core Difference Between GEO and SEO **SEO buys a ranked position you can hold; GEO buys a mention inside an answer that gets rewritten every prompt, and the two pay off in different currencies.** SEO's return is a click from a spot you defend. GEO's return is a citation in a probabilistic answer with no position to hold, so you fund a share of many answers rather than one rank. That is why this is a budget question, not a swap. Rankings and citations come apart in practice: a page can sit outside Google's top 10 and still get quoted by ChatGPT, and a page at position 1 can be left out of the answer entirely. If you want the wider acronym picture, including where answer engine optimization (AEO) fits, our [GEO vs AEO vs SEO breakdown](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) is the triage page. This one stays on the two disciplines most budgets actually trade off: search and AI, where GEO means [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), the work of getting cited by AI answer engines. ## Why GEO and SEO Share a Foundation Here is the part that decides how you should spend: **most AI answer engines do not run their own web-scale crawl. They lean on a search index to find candidate pages, then fetch those specific pages to read and quote them.** Ask ChatGPT or Google's AI features a question and the system rewrites it into several queries, pulls candidates through search-style retrieval (Bing plus OpenAI's own index for ChatGPT, Google's own for AI Overviews), and writes an answer from what comes back. Perplexity runs more of its own retrieval, but the pattern holds: to be quoted, you first have to be findable. That retrieval step is [query fan-out](https://geotoolbox.ai/blog/query-fan-out), and it means the SEO foundation, being crawlable, indexable, and credible, is the layer GEO is built on. [How AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) walks the full loop. The data backs the dependency. A [Writesonic study of over a million AI Overviews](https://writesonic.com/blog/ai-citations-from-serp-results-study) found 40.58% of AI citations came from pages in Google's top 10 for the query. The top 10 is still the single biggest door into an AI answer, even though the majority of citations now come from beyond it, often from a page that won a fan-out sub-query no rank tracker shows you. Note what that implies: you do not need to rank first for the head term to be cited, you need to rank for one of the sub-questions the answer is built from. Google says the same thing from the other side: its [AI features run on core Search ranking and indexing](https://developers.google.com/search/docs/appearance/ai-features), with no special markup that turns them on. So the foundation is shared. What differs is the slice on top, and that slice is where the budget question lives. ## GEO vs SEO, Side by Side The deltas that matter day to day are narrow but real. SEO optimizes for a crawler that ranks; GEO optimizes for a model that retrieves and rewrites. That changes what you measure, who you compete with, and which content actually wins.
DimensionSEOGEO
Unit of successA ranked position and its clickA citation inside an AI answer
What gets pickedPages, ranked by relevance and authorityPassages, selected by extractability and trust
Where you appearThe results listInside a written, sourced answer
How it is measuredRankings, clicks, sessionsShare of voice, mention rate, citation rate
Content that winsDepth, links, keyword coverageClear claims, quotes, stats, consistent entity
Visibility patternA position you holdA share that shifts prompt to prompt
The last two rows are where the money splits. Rankings still reward links and depth; citations reward being a clearly defined entity, named and described the same way across the web, so [entity SEO](https://geotoolbox.ai/blog/entity-seo) and digital PR earn more of the GEO budget than another link-building push. Because the answer reshuffles every run, you fund a share of voice rather than a rank you hold.
![GEO vs SEO scorecard showing which discipline leans for each job, with a shared foundation.](/blog/geo-vs-seo/geo-vs-seo-scorecard.png)
SEO and GEO lean different ways by job, but sit on one shared foundation. Neither is an overall winner.
## When SEO Still Wins GEO is additive, not a replacement, and for whole classes of query and business SEO is still the better place to spend. **Most searches still have no AI answer to be cited in.** Per [Pew Research's July 2025 study](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) of 68,879 searches, an AI summary appeared on about 18% of them. Where one does fire it suppresses the click hard, with Pew putting the click rate at 8% with a summary versus 15% without, but that cuts both ways: at Pew's mid-2025 measurement most searches carried no summary to be cited in. That share has shrunk since, yet still covers the majority of queries, and on them the ranked link is the whole game. High-intent and YMYL queries still click. When someone is ready to buy, book, or reach a specific site, they act on a result rather than ask a chatbot to do it for them. It also holds for high-stakes medical, legal, and financial decisions, where people click through to verify an accountable source rather than take a summary at its word. These reward a strong ranked page and a fast path to act, not a citation in someone else's answer. **AI tends to get the research question, search gets the buying one.** People reach for an assistant to explore and compare, then return to search to transact. a16z, [citing Semrush, puts AI prompts at about 23 words versus 4 for a traditional search](https://a16z.com/geo-over-seo/); treat the exact number as directional, since Semrush's own clickstream reports shorter prompts, but the direction fits the behavior, longer exploratory questions upstream and decisive clicks downstream. The bottom of the funnel still lives in search. **Being findable comes before being cited.** Because engines retrieve from a search index, a page that is not crawlable or indexable rarely gets quoted. A low-authority site can still win a narrow sub-query it answers well, but it cannot skip the groundwork. If you are early, the fastest route to AI visibility is the SEO foundation: crawlable, credible, findable first. ## How to Split Your Budget Between SEO and GEO There is no universal percentage. SEO and GEO are interdependent: the same work that lifts a ranking often lifts a citation, so a clean "X% to GEO" split misrepresents how the money actually works. The vendor guidance that circulates varies widely, from an even split to the large majority of effort on SEO depending on who you ask, which tells you it is a judgment call rather than a formula. Start from your business, not a slider. This is where the lean sits by model:
If you areLeanWhyFirst GEO dollar
Local serviceSEOBuyers want one nearby result and a click to actReachability + a clean profile
EcommerceSEO-leanCheckout stays on-site, but product discovery is moving into AI, so feed it tooStructured product data
B2B or SaaSToward GEOLong research cycles run through AI comparisons; the win is being the recommended brandThird-party citations and entity trust
Publisher or mediaSEOAd models need the visits, but publishers are the most exposed to summarizingTrack citation leakage first
New or low-authoritySEO foundationCitations follow discoverability you have not built yetNothing until you can be found
If you need a place to start, most programs sit somewhere between 85/15 and 70/30 in favor of SEO, tilting toward 60/40 for high-consideration B2B where buyers lean on AI to build a shortlist. Treat that as a dial, not a rule: begin near 80/20 and let your own exposure move it. Then size the GEO slice to that exposure, not a benchmark. Two numbers set it: how many of your money queries actually trigger an AI summary (the 18% average runs higher in some categories), and how much of your visibility already leaks to zero-click answers in Search Console. AI answers have only spread since early 2025, when Semrush measured [AI Overviews on 6.49% of US desktop searches in January 2025](https://www.semrush.com/blog/semrush-ai-overviews-study/) across more than 10 million keywords, a share that more than doubled over the following two months, so revisit the split as your category gets more coverage. In practice, hold your total search-and-content budget steady and fund GEO by redirecting part of your existing content, PR, and link spend rather than adding a new line, letting measurement justify every shift after that. You also do not need to buy a separate service to do it. Most of the work is your own SEO and PR redirected, a point we make in [GEO services vs software](https://geotoolbox.ai/blog/geo-services-vs-software). ## How to Migrate an SEO Program to GEO Most guides hand you a GEO checklist. What an existing SEO team actually needs is a sequence: how to evolve the program you already run into one that earns citations too. Four phases, in order. 1. **Fix reachability first.** The crawlers that fetch pages for live answers, OAI-SearchBot and ChatGPT-User, PerplexityBot and Perplexity-User, Claude-User, mostly read raw HTML and do not run JavaScript. A robots.txt rule, a bot-management block, or client-side rendering can quietly lock them out even when Googlebot gets in. (Blocking the training-only crawlers like GPTBot keeps you out of model training but does not stop you being cited, so leave the retrieval ones in.) This is the cheapest win and usually a same-day change. Confirm the retrieval crawlers can reach your pages with our [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) before anything else. Nothing downstream matters until this passes. 2. **Redesign the content for extraction.** Shift from pages that rank to passages a model can lift. Lead each section with a direct 40-to-60-word answer, phrase headings as the questions people ask, and put evidence in the body. The [GEO research paper that coined the term](https://arxiv.org/abs/2311.09735) found that adding citations, quotations, and statistics raised a page's visibility in generated answers by up to 40%. That is not a new skill, it is [AI content optimization](https://geotoolbox.ai/blog/ai-content-optimization) applied to a new reader. 3. **Build off-site entity trust.** Rankings reward links; citations reward a consistent, well-described entity. Move some link budget into digital PR and getting named accurately across the sources models lean on, from Wikipedia and Wikidata to Crunchbase, LinkedIn, and the main directories in your category. The goal is that your brand is described the same way wherever a model looks. 4. **Evolve your KPIs.** Add citation rate, share of voice, and mention rate alongside rankings and sessions, and expect engines to disagree with each other. The [full AI-search playbook](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) covers the tactics under each phase; the shift here is what you hold yourself accountable to. ## The Measurement Problem Nobody Warns You About Budgets stall on GEO for a plain reason: you cannot manage what you cannot see, and AI visibility is genuinely harder to measure than a ranking. Three things make it slippery. Answers are probabilistic. SparkToro ran the same prompts thousands of times and found [AI engines returned an identical list of recommended brands under 1% of the time](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/), and the same order closer to 1 in 1,000. Any score you get samples that noise, so treat it as a trend line, not a census. Citations do not always name you. Semrush's [Ghost Citations study found 62% of AI citations](https://www.semrush.com/blog/the-ghost-citations-study/) reference a source without naming the brand, so a mention-only tracker undercounts your real influence. And the visits that do come through often arrive with no referrer and land in analytics as direct traffic, which undercounts again. Attribution will not be clean, but you can close some of the gap. Watch branded-search and direct-traffic lift in the weeks after you gain citations, add a "how did you hear about us" prompt at signup, and treat share of voice as the leading indicator while conversions catch up. Set expectations while you are at it: most AI citations behave like an impression, not a click. GEO's near-term return is upper-funnel, being the brand the answer names, so budget it as brand influence rather than direct traffic. The fix is to track the right things: share of voice across engines, mention rate, citation rate, and where you land inside an answer. In the sites we scan, a page is often invisible in AI answers for a reason a visibility check surfaces quickly, so measure before you optimize. Our guides to [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) and [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) cover the metrics in depth. ## Is GEO Just SEO With a New Name? Partly, and that is worth conceding. Ask working SEOs and many will say most GEO advice is plain good SEO with a fresh label, and [threads full of that skepticism](https://www.reddit.com/r/SEO/comments/1o0ta17/whats_the_real_difference_between_approaching_seo/) are not wrong: reachable pages, clear answers, real authority, and consistent branding earn both rankings and citations. If a "GEO package" is reselling you those fundamentals at a new price point, you are paying twice for one program. The sharper objection is about efficacy, not price. In several practitioner write-ups, the accounts where GEO "worked" were the ones already ranking well, which suggests the citations were reflecting existing SEO rather than adding to it. It is a fair challenge. The incremental layer, extractable formatting and entity trust, is worth doing, but its standalone return is still thinly evidenced. Which is exactly why you measure before you scale it. But the rebrand framing also hides a genuinely new layer. Retrieval changes which pages get surfaced, citation changes how credit is assigned, and measurement changes what you can even see, none of which existed when the job was purely ranking. The practical takeaway is not "GEO is fake" or "SEO is dead." It is that the fundamentals now serve two surfaces, and the new work is the thin slice on top plus the tracking to prove it moved. Spend on that slice, not on a second invoice for the base. ## Frequently Asked Questions ### Is GEO replacing SEO? No. GEO adds a surface, it does not remove one. Most searches still return no AI answer to be cited in, and AI engines mostly retrieve candidates through search-style discovery, so the authority that makes you findable is what makes citations possible. Treat GEO as an extension of SEO, not a successor. ### Is SEO dead in 2026? No, though the click is under pressure. AI answers resolve more queries without a visit, so some sites see organic traffic slide. The response is to extend SEO toward citations and answers, not abandon it, because the fundamentals that earn a ranking are the same ones that earn a citation. We walk through the data behind that in [is SEO dead?](https://geotoolbox.ai/blog/is-seo-dead). ### Is GEO the same as geo-targeting or local SEO? No, and the name collision causes real confusion. GEO here means generative engine optimization, getting cited by AI answer engines like ChatGPT and Perplexity. It has nothing to do with geographic targeting, geolocation, or local search. ### How much of my budget should go to GEO? There is no universal number. Hold your total budget steady and fund GEO by redirecting part of your existing content, PR, and link spend, sized to how many of your money queries trigger AI answers and how much visibility you already lose to zero-click. Let tracking justify each shift rather than committing to a fixed split up front. ### Can I do GEO without doing SEO first? Rarely well. Because AI engines mostly find candidates through search-style discovery, a page that cannot be found or indexed usually cannot be cited. New or low-authority sites get the fastest AI-visibility gains by building the crawlable, credible SEO foundation first. ### Can optimizing for GEO hurt my SEO? Rarely, and mostly they pull the same way. The one real tension is over-trimming a page into extractable snippets at the expense of the depth that earns rankings. Keep the substance, add the answer-first structure on top, and the two reinforce each other. ### Is GEO the same as AEO? They overlap but are not identical. Answer engine optimization targets the answer slot on search surfaces, while GEO targets citations inside generated, conversational answers. Our [3-way breakdown](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) untangles all three. ## Where to Start For most businesses the verdict is the same: keep SEO as the foundation and layer GEO on top, sized to your actual AI exposure. Do not defund the work that still brings the traffic to chase a surface that, for now, sends a fraction of it: AI referrals were about [0.32% of all website traffic in SE Ranking's June 2026 study](https://seranking.com/blog/ai-traffic-research-study/), growing fast but still tiny next to search. The two share a base, so the smart move is to run the fundamentals once and spend the incremental effort where your own data says you are losing visibility. That last part is the catch. You cannot split a budget you cannot measure, and AI visibility does not show up in a rank tracker. Before you move a dollar toward GEO, get a read on where you actually stand across the AI answers your buyers see. geotoolbox tracks your share of voice and citations across ChatGPT, Perplexity, Google AI Overviews, and more, so the split is a decision backed by numbers rather than a guess. Start with a [look at your AI visibility](https://geotoolbox.ai/features/domain-overview), then let what it shows decide how much of the budget shifts. ## Sources - AI Citations From SERP Results Study - Writesonic - `writesonic.com/blog/ai-citations-from-serp-results-study` - Google Users Are Less Likely to Click When an AI Summary Appears - Pew Research Center, July 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` - GEO: Generative Engine Optimization (paper) - Aggarwal et al., arXiv 2311.09735 - `arxiv.org/abs/2311.09735` - AIs Are Highly Inconsistent With Their Recommendations - SparkToro - `sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility` - Semrush AI Overviews Study - Semrush - `semrush.com/blog/semrush-ai-overviews-study` - The Ghost Citations Study - Semrush - `semrush.com/blog/the-ghost-citations-study` - GEO Over SEO - Andreessen Horowitz (a16z) - `a16z.com/geo-over-seo` - AI Features and Your Website - Google Search Central - `developers.google.com/search/docs/appearance/ai-features` - What's the Real Difference Between SEO and GEO (discussion) - r/SEO - `reddit.com/r/SEO/comments/1o0ta17` - AI Traffic Research Study - SE Ranking, June 2026 - `seranking.com/blog/ai-traffic-research-study` --- ## Grok 4.6: Specs, Benchmarks, Real Cost & the Honest Verdict (2026) > Grok 4.6 is xAI's agent-first model, out August 12, 2026. A close read of the specs, the real benchmark table, the 200K price cliff, and AI-search impact. - Canonical: https://geotoolbox.ai/blog/grok-4-6 - Published: 2026-08-13 · Updated: 2026-08-13 Grok 4.6 is xAI's newest model, and the day after it launched we ran a quick test: we asked two frontier AI assistants, with web search switched off, what Grok 4.6 was. One, running on training data that ends in early 2024, knew Grok only up to version 1.5. The other, current to late 2025, stopped at Grok 4.1. Neither knew 4.6 existed. Only an assistant that searches the live web can see it at all, and that gap is a small preview of why this release matters well beyond the benchmark charts. This is the plain, current version of what Grok 4.6 is: what xAI shipped on August 12, 2026, what it really costs once you read past the headline price, how the benchmark table looks when you read the losses first, and what a model built for agents means for whether AI can find your site. The numbers here come from xAI's own announcement and model card, cross-checked against Artificial Analysis.
![Three AI engines and the latest Grok each knows: a GPT-family model with an early-2024 cutoff knows only up to Grok 1.5, a Claude-family model current to late 2025 stops at Grok 4.1, and only live web retrieval sees Grok 4.6, released August 12 2026.](/blog/grok-4-6/grok-4-6-engine-knowledge.png)
Checked the day after launch: only an engine that retrieves the live web can see a model released yesterday.
## What Is Grok 4.6? **Grok 4.6 is xAI's frontier AI model, released on August 12, 2026, and built for long-running agents, coding, and knowledge work.** It is the newest version of [Grok](https://geotoolbox.ai/blog/what-is-grok), the [large language model](https://geotoolbox.ai/glossary/large-language-model) assistant made by xAI, the company Elon Musk founded in 2023 and which now operates as SpaceXAI after the 2026 SpaceX merger. If you have been tracking the version numbers, 4.6 lands about a month after [Grok 4.5](https://geotoolbox.ai/blog/grok-4-5) and slots in as the current flagship. Grok 4.6 is not a bigger model than 4.5. xAI describes a longer round of training on top of 4.5's foundation rather than a larger base, and secondary coverage reads the underlying model as unchanged in size. That matters because it changes how you should read the benchmark numbers below. The model is aimed at people who build software and run agents, not at people asking a chatbot for trivia, and at launch it shipped for developers and inside apps, not in xAI's own consumer chatbot. Grok 4.6 reads text and images, writes text, holds a 500,000-token context window, and adds a new "xhigh" reasoning setting above the levels 4.5 offered. Its knowledge cutoff is February 1, 2026 (the model card gives a January 2026 pretraining cutoff), so anything newer than that, including its own launch, it only knows through live search tools, not from memory. ## Is Grok 4.6 Out? What xAI Actually Shipped Yes. Grok 4.6 went live on August 12, 2026, announced by SpaceXAI (the name xAI now trades under). Musk had publicly floated a rollout "around August 7," so the date slipped a few days, and the model shipped with a full evaluation table, a model card, and live API access rather than a teaser. What actually changed is narrower than a new version number suggests, and xAI is fairly open about it. According to [xAI's announcement](https://x.ai/news/grok-4-6), Grok 4.6 came from a longer supplemental training run than 4.5 got, regenerated fine-tuning data, and reinforcement learning in agentic environments. xAI does not itself claim a larger base model; the "same ~1.5-trillion-scale base, held constant" reading comes from secondary coverage. Either way, the work went into the training layers, not raw scale. A few things are genuinely new. The first is the agent focus: xAI says the model stays on a task across many steps without drifting, and that on longer runs it started checking its own work before moving on. That is vendor self-observation, not an independently measured result, so test it against your own workload. The second is the new **xhigh** reasoning-effort level, which sits above the low, medium, and high settings 4.5 shipped with. The third is the coding tuning: Grok 4.6 was developed in collaboration with Cursor and received supplemental training on anonymized Cursor workflow data, building on the Cursor collaboration behind 4.5. Two claims worth pinning down, because the coverage blurs them. Some write-ups described the Cursor relationship as xAI acquiring or investing in Cursor; xAI's own materials call it a training collaboration, so leave the ownership version alone until a primary source confirms it. And there is no open-weights release: Grok 4.6 is a closed, hosted model with no self-hosting path. ## Grok 4.6 Specs at a Glance Here is the spec sheet, from [xAI's developer docs and model card](https://docs.x.ai/developers/grok-4-6).
PropertyGrok 4.6
Model IDgrok-4.6
ReleasedAugust 12, 2026
Context window500,000 tokens (unchanged from Grok 4.5)
Knowledge cutoffFebruary 1, 2026
ModalitiesText and image input; text-only output; no output limit
Reasoning levelsLow, medium, high (default), xhigh (new)
ToolsFunction calling, web search, X search, code execution
Price$2 / $0.50 / $6 per 1M tokens (input / cached / output) below 200K prompt tokens; doubles above
AvailabilityxAI API, Grok Build (default), Cursor (all plans), Microsoft Office add-ins; OpenRouter, Vercel, Cloudflare, and other gateways
Open weightsNo; no self-hosting
Two details in that table get missed in most write-ups. The context window did **not** grow; it was already 500K on Grok 4.5, so anyone implying a bump is wrong. And the price has a second half that only appears once a request's prompt gets past the 200K mark, which we break down next because it lands squarely on the agent workloads this model is sold for. ## Grok 4.6 Benchmarks: Read the Losses First Grok 4.6 scores **61 on the Artificial Analysis Intelligence Index**, a composite of nine benchmarks, up five points from Grok 4.5's 56 and level with GPT-5.6 Sol Max. That figure is third-party: [Artificial Analysis](https://artificialanalysis.ai/models/grok-4-6) runs the index. [VentureBeat](https://venturebeat.com/technology/spacexai-debuts-grok-4-6-overtaking-kimi-k3s-performance-and-matching-gpt-5-6-sol-for-worlds-third-best-on-artificial-analysis) called it the world's third-best model, behind Anthropic's Claude Opus 5 and Fable 5. Read that carefully: on Artificial Analysis's own page the Grok 4.6 (high) variant sits sixth of the 184 model variants it tracks as of mid-August 2026, because "third best" counts flagship models and the two ahead of it are both Anthropic's. The more useful exercise is to read xAI's own evaluation table. Here is the full table xAI published, with the best score in each row in bold.
BenchmarkGrok 4.6 (High)Grok 4.5 (High)GPT-5.6 Sol MaxFable 5 Max
AA Intelligence Index61566162
GDPVal-AA v21753152617281741
CursorBench v3.269.9%66.7%67.2%70.5%
DeepSWE v1.165.9%54%73%70%
FrontierCode v1.1 (Extended)61.3%56.6%60.6%63.6%
APEX-Agents57.5%47.1%56.7%59.2%
Terminal-Bench v3.026%15.7%34.6%34.1%
APEX-SWE56.4%53.6%n/a58.8%
AA-Briefcase1577131315021574
Harvey LAB15.8%12.9%2.5%11.3%
Read down the bold column: **Fable 5 Max wins more rows than anyone else, five of the ten.** GPT-5.6 takes two, DeepSWE and Terminal-Bench. Grok 4.6 posts the top raw score in three: GDPVal-AA v2, AA-Briefcase, and Harvey LAB. Two of those come with an asterisk. Its GDPVal and AA-Briefcase margins over Fable are only a handful of Elo points (1753 to 1741, and 1577 to 1574), which MarkTechPost reads as inside the noise, and both are wins only because this four-model table leaves out Claude Opus 5, which tops the Intelligence Index overall and, by xAI's own model card, leads both of those evals. That leaves Harvey LAB as Grok's one clear win, and it too deserves a caveat: GPT-5.6 scores just 2.5% there against Grok's 15.8%, a gap wide enough to look like a scoring or harness artifact until someone reproduces it. Head to head with GPT-5.6, the coding rows split evenly: Grok 4.6 edges CursorBench and FrontierCode, GPT-5.6 leads DeepSWE and Terminal-Bench. One more thing to read carefully: this is xAI's table, not a neutral one. xAI chose which benchmarks appear, and each rival's column is that rival's best self-reported or publicly available result rather than a controlled head-to-head. The benchmarks themselves are third-party suites, not xAI in-house tests: Artificial Analysis (the Intelligence Index, GDPVal-AA v2, and AA-Briefcase), Vals AI (Harvey LAB), Cursor (CursorBench), Mercor (the APEX evals), Cognition (FrontierCode), Datacurve (DeepSWE), and the Harbor team at Stanford with the Laude Institute (Terminal-Bench). The model card credits Artificial Analysis, Mercor, and Datacurve with running the GDPVal, APEX-SWE, and DeepSWE results themselves; the rest xAI appears to have run against those external suites. Where the vendor's hand shows is the selection: Claude Opus 5 has no column, and CursorBench is scored on Cursor's own harness by the same partner whose anonymized workflow data went into training this model. So read it as a vendor-curated scorecard rather than a controlled head-to-head: Grok 4.6 is a clear step up from 4.5, level with GPT-5.6 on the broad independent index, and behind the top Claude models. ## What Grok 4.6 Really Costs The headline is cheap: **$2 per million input tokens and $6 per million output**, the same as Grok 4.5 and, by xAI's own framing, about half what the flagship Claude and GPT models charge. That is the number every roundup leads with. It is also only half the story. For how the rates line up across Grok versions, see our [Grok pricing guide](https://geotoolbox.ai/blog/grok-pricing). Send a request whose prompt hits the **200K mark and the rate doubles to $4 in, $12 out, and $1 for cached input, applied to every token in that request**, not just the tokens past the line. Below that line, cached input is $0.50 per million, but that rate is not automatic: xAI's docs tell you to set a `prompt_cache_key` (or the `x-grok-conv-id` header on Chat Completions), or your requests scatter across servers and you pay the full $2 input rate on a cache-cold one. The awkward part is where the cliff sits: 200K is exactly the context a long-running agent working across a large codebase or a 500K-token corpus will blow through, which is the workload xAI markets this model for. The 500K window is real, but the second half of it is billed at double. For long agent loops, xAI points to context compaction as the way to stay under the line. The per-token price is only part of the cost, and the fuller picture cuts both ways. On the plus side, Artificial Analysis's own runs show Grok 4.6 finishing its AA-Briefcase agent workload in roughly 53 turns and about half a billion input tokens, against about 103 turns and two billion for Claude Opus 5 Max, a large per-task efficiency edge (VentureBeat's caveat: whether it carries from those controlled runs into production is untested). On the minus side, its cost-per-completed-task lands mid-pack on the same site, less economical than several cheaper models and than Grok 4.5 itself, with throughput near the median. Put together: cheap per token, often fewer turns per task, but mid-pack once you price the whole workflow, and double above 200K. There is also a "fast" variant at double the standard price. xAI has not published a separate model ID, latency figure, or throughput spec for it, so at launch you get the price without the specification. For the first week, Grok Build and Cursor include 2x usage, meaning double the usual allowance while you try it, which does not change the underlying rates. ## Grok 4.6 vs Grok 4.5, GPT-5.6 and Claude If you are deciding between models, the short version is that Grok 4.6 is the agentic upgrade over 4.5 and a mid-priced alternative to the frontier leaders, not the top model on the independent index.
QuestionThe short answer
Is 4.6 worth it over Grok 4.5?Yes for multi-step agentic and coding work (+5 on the intelligence index, large jumps on the agent benchmarks). For scoped, validated, cache-heavy prompts, 4.5 is still fine; it lists at the same price, so the case for staying is simplicity, not a lower rate.
vs GPT-5.6 Sol Max?Level on the composite index (61 each). Head to head, the coding rows split two apiece: Grok edges CursorBench and FrontierCode, GPT-5.6 leads DeepSWE and Terminal-Bench. Roughly a wash, decided by price and your stack.
vs Claude Opus 5 / Fable 5?Both Claude models sit above Grok 4.6 on the index: Fable 5 wins more rows than anyone in xAI's own chart, and Opus 5 tops the index overall. Grok's argument is cost, not raw capability. See our Grok vs Claude breakdown.
vs Kimi K3?Grok 4.6 passes Kimi K3 on the intelligence index. Kimi's pitch is open weights; Grok's is a managed, tool-rich API.
Whatever the current leaderboard says, these standings turn over every few weeks now. A ranking that is true in August can flip by October, so timestamp any "best model" claim you rely on. For how the two most common rivals stack up in practice, see [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt). ## How to Access Grok 4.6 At launch, Grok 4.6 is mostly a developer product. Per xAI's model card, it is available through the xAI API (as `grok-4.6`), as the default model in Grok Build, in Cursor on every plan tier, and, notably, as the default model in the Microsoft Word, PowerPoint, and Excel add-ins, a direct move onto Copilot's turf. It is also routable through OpenRouter, Vercel, Cloudflare, and other gateways. What it is **not**, yet, is a consumer chatbot. The model card says xAI plans to add Grok 4.6 to its consumer surfaces (the web app, mobile apps, and Grok-in-X) "at a later date," so if you are asking whether Grok 4.6 is the model answering you at grok.com or in the X app today, the answer is probably not yet. Is it free? xAI's launch page advertises a "Try it in Grok Build for free" option, while press coverage puts Grok Build behind the $30-a-month SuperGrok plan, so check x.ai/build before assuming free access. API use is paid at the rates above. The clearest free element is the first-week 2x usage promotion in Cursor and Grok Build. If you are in the European Union, verify availability before you plan around it. Grok 4.5 was blocked across all 27 EU states until mid-July after its launch while xAI completed EU AI Act evaluations, and no source confirms Grok 4.6's day-one EU status either way, so treat it as an open question rather than an assumed yes. ## Grok 4.6, Agents, and Whether AI Can Find You This part matters even if you never call the API. Grok 4.6's headline feature is agents: models that run for many steps, use tools, and pull information as they go. Two of its four built-in tools are web search and X search, and its own knowledge stops at February 1, 2026. That combination means the model does not "know" the current web; it retrieves it, live, every time a question runs past its cutoff. It could not tell you about its own launch without searching for it. That is the whole argument for [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) in one model. When an agent answers a question about your company, your product, or your market, it is reading whatever it can fetch and parse in that moment. If your pages are slow, blocked to its crawler, or structured so the important facts are buried in scripts and images, the agent works from whatever it found instead, which is often a competitor or a forum thread. The model getting better at agentic work raises the stakes on being retrievable, because more of what people learn about you now passes through a machine reading your site on their behalf. This is the lane we work in. Grok, like the other engines, reaches your content through a bot, and the first question is whether that bot can even load your pages. You can check which AI crawlers reach your site, and whether they are being served or blocked, with our [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker), and get a broader read on whether your pages are structured to survive an agentic pass with the [AI Readiness](https://geotoolbox.ai/tools/ai-readiness) check. A new model version is a good prompt to confirm the basics: the fastest capability gain for you is not the model's, it is making sure the pages it retrieves are yours. For the mechanics of how these bots crawl and render, see our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers). ## Should You Use Grok 4.6? The Verdict Grok 4.6 is a real step up from 4.5 and a credible mid-priced frontier model, though not the top model on the independent index. Whether it is right for you comes down to a few concrete questions. Use it if you are doing multi-step agentic or coding work, want tool-rich API access without managing weights, and your requests stay under 200K prompt tokens, where the price is competitive. Think twice if your workload routinely crosses that 200K line, since the doubled rate erases the cost advantage on exactly the long-context agent runs the model is pitched for. Stay on Grok 4.5 for scoped, validated, cache-heavy prompts where the extra capability does not pay for itself. Two more factors belong in the decision. On governance, xAI's model card documents autonomous behavior in detail but includes little formal autonomy-risk evaluation, which is worth weighing for a regulated or enterprise deployment; and the wider picture matters too, since X and the earlier xAI organization face open UK and EU investigations (Ofcom, the ICO, and a formal EU DSA probe) that concern the platform's history rather than any finding against the Grok 4.6 API. On timing, Musk has signaled on X that a larger Grok 4.7 is "weeks" away, reportedly a genuine scale jump rather than another post-training pass, and his dates have slipped before, so treat it as a target rather than a plan. For what is known about the next jumps, see [what we know about Grok 5](https://geotoolbox.ai/blog/grok-5). ## Frequently Asked Questions ### Is Grok 4.6 free? Not as a consumer chatbot yet. At launch it ran on developer surfaces (the API, Cursor, Grok Build, and the Office add-ins), and xAI's model card says consumer access on the web, mobile, and X is coming "at a later date." xAI's launch page advertises a free Grok Build trial, though coverage puts Grok Build inside the $30-a-month SuperGrok plan, so check x.ai/build before assuming free access. API use is paid. ### How much does Grok 4.6 cost? $2 per million input tokens and $6 per million output, with cached input at $0.50, below 200K prompt tokens. Above that threshold every token in the request is billed at double: $4 input, $12 output, $1 cached. A separate "fast" variant costs twice the standard rate. ### What is Grok 4.6's context window? 500,000 tokens, the same as Grok 4.5. It did not increase. Note that the doubled pricing kicks in at 200K prompt tokens, well before the window is full. ### Is Grok 4.6 better than Grok 4.5? On the benchmarks, yes: it scores 61 on the Artificial Analysis Intelligence Index versus 56 for 4.5, with sizable gains on agent and coding tests. In practice it is the better choice for multi-step agentic work, while 4.5 remains reasonable for focused, cache-heavy prompts. The two list at the same price, so the reason to stay on 4.5 is simplicity, not cost. ### Is Grok 4.6 better than GPT-5.6 or Claude? It ties GPT-5.6 Sol Max at 61 on the composite index, and against GPT-5.6 the coding benchmarks split evenly. Anthropic's Claude Opus 5 and Fable 5 both sit above Grok 4.6 on the index. Grok's advantage is price, not top-end capability. ### How many parameters does Grok 4.6 have? xAI has not published an exact parameter count. The model card describes it as part of a "1.5-trillion-scale" family, so read "1.5T" as a family-level label rather than a confirmed spec. ### When is Grok 4.7 or Grok 5 coming? Musk has said on X that Grok 4.7 is "weeks" away at a larger scale; his dates have slipped before, so his timelines are targets, not commitments. No primary source gives a Grok 5 date, so treat any specific timeline for it as speculation. ## Sources - Introducing Grok 4.6 - xAI (SpaceXAI), Aug 12, 2026 - `x.ai/news/grok-4-6` - Grok 4.6 model overview and specifications - xAI Docs - `docs.x.ai/developers/grok-4-6` - Grok 4.6 Intelligence Index, cost and speed - Artificial Analysis - `artificialanalysis.ai/models/grok-4-6` - SpaceXAI debuts Grok 4.6 - VentureBeat, Aug 12, 2026 - `venturebeat.com/technology/spacexai-debuts-grok-4-6-overtaking-kimi-k3s-performance-and-matching-gpt-5-6-sol-for-worlds-third-best-on-artificial-analysis` - SpaceXAI Releases Grok 4.6 - MarkTechPost, Aug 12, 2026 - `marktechpost.com/2026/08/12/spacexai-releases-grok-4-6` - xAI's Grok 4.6 Holds the Base Model Constant - FourWeekMBA, Aug 12, 2026 - `fourweekmba.com/ai-xai-grok-4-6-post-training-turn` --- ## Semrush Review (2026): We Tested It on Three Real Sites > A hands-on Semrush review from a paying account: accuracy tested against Search Console on three sites, 2026 pricing decoded, and the complaints that matter. - Canonical: https://geotoolbox.ai/blog/semrush-review - Published: 2026-08-02 · Updated: 2026-08-11 Plenty of Semrush reviews are feature tours written from a quick trial. This one is written from the paid account we use for client work every day, checked against Google Search Console on three real sites, and updated for what Semrush became in 2026: an Adobe-owned platform selling AI visibility alongside classic SEO. Two things to know before reading: Semrush links here are affiliate links, and we build an AI visibility tool of our own, which makes part of what we review below a competing product. ## Semrush Review: The Short Answer **Semrush is the strongest all-in-one SEO platform you can buy in 2026, and the most annoying one to pay for. Verdict: 4.25/5, docked mostly for the billing traps and seat math rather than the product.** The core toolkit earns its reputation. Keyword Magic is the best keyword research tool we have used, the paid-search intelligence has no real equivalent in [Ahrefs](https://geotoolbox.ai/blog/semrush-vs-ahrefs), and the reporting features are what justify the invoice for agencies. The hard part is everything around the product. The price on the website is not the price you end up paying once seats and add-ons stack up. The AI Visibility Toolkit's advertised $99 a month carries a "billed annually" line on the tab the pricing page opens to, but the monthly tab offers the same $99 with no annual discount at all. And billing and cancellation dominate the community complaints we monitored, ahead of any product flaw. Here is the short version: **Buy it if** you run an agency or an in-house team that will use three or more of its toolkits, you need PPC competitive data, or you want AI search visibility and classic SEO in one subscription. **Skip it if** you are a solo operator with one small site. You would be paying for dozens of tools to use five, and the cheaper tiers fence off the features that make the price defensible. Every number below is either a figure from our own dashboards checked against Search Console, or a price verified on Semrush's own pages this week. We label which is which.
![Semrush Domain Overview for geotoolbox.ai in August 2026, showing the new AI Search panel with 25 cited pages next to the classic SEO panel with Authority Score 11, 1.7K organic traffic, and 1.6K organic keywords.](/blog/semrush-review/semrush-domain-overview-2026.jpg)
The 2026 Domain Overview: an AI Search panel now sits beside the classic SEO metrics. This is our own domain.
## How We Tested It This review is based on the Semrush account we pay for and use daily, not a trial spun up for a screenshot. We ran Semrush against three sites we have full data access to: our own site (a small, young domain), a mid-size manufacturer site, and a large marketplace in the high five figures of monthly clicks (the latter two are client sites, anonymized). For each one we compared Semrush's numbers with Google Search Console, with Ahrefs (which we also pay for), and, for the AI-citation check, with Bing Webmaster Tools. We pulled the same dashboards twice, twelve days apart, to see how the numbers move. That methodology matters because it separates two questions reviews usually blur: what Semrush can do, and whether its numbers are right. A feature tour cannot answer the second question. Ground truth can. We published the first round of that data in our [Semrush vs Ahrefs comparison](https://geotoolbox.ai/blog/semrush-vs-ahrefs) in July. This review extends it with the full 2026 pricing picture, the AI Visibility Toolkit, and the buying decisions the comparison did not cover. ## What Semrush Is Best At Strip away the marketing and Semrush is really three exceptional products wearing one login. ### Keyword Magic and Keyword Research The Keyword Magic Tool is the closest thing to a consensus in SEO tooling: even the community threads that trash Semrush's billing concede this one. Type a seed term and it returns keyword variations grouped into semantic clusters, with intent labels, difficulty scores, CPC, and SERP feature flags attached to each row. Two things make it better than the equivalents. The clustering does hours of manual sorting for you, and the personalized difficulty option estimates how hard a keyword would be for **your** domain rather than in the abstract. We use it weekly for keyword research on client campaigns, and in that use the variation coverage runs broader than what Ahrefs returns for the same seeds, which matches the database gap both vendors advertise. We have not run a formal seed-by-seed count, so read that as working impression rather than a measured result.
![Semrush Keyword Magic Tool results for the seed 'ai visibility': 16,483 keywords totaling 311,800 monthly searches, grouped into clusters like tool, search, and brand, each row showing intent, volume, keyword difficulty, and CPC.](/blog/semrush-review/semrush-keyword-magic-tool-2026.jpg)
One seed term, 16,483 keywords, from our account in August 2026. The clustering is what replaces the manual sort.
### Competitive and PPC Intelligence This is the moat. Organic Research shows a competitor's ranking pages and estimated traffic split, the Keyword Gap and Backlink Gap tools map exactly where rivals rank and earn links that you do not, and Advertising Research exposes their paid keywords, ad copy history, and where the budget goes. Ahrefs carries some paid-search data too, but nowhere near this depth: no ad copy history, no comparable spend view. Standalone PPC spy tools cost real money on their own. If you sell to clients, Advertising Research alone can carry a pitch meeting. It's the feature we point to when someone asks why the subscription survives our annual tool cull. ### Site Audit and Agency Reporting Site Audit's crawl budget scales with plan tier and, more usefully, it ranks what it finds by impact instead of dumping an error list. Each issue carries a plain-language explanation of why it matters and how to fix it. For agencies, the white-label scheduled PDF reports are the quiet workhorse (note that on current plans some branding and scheduling options sit in a low-cost reports add-on, $10 to $20 a month, rather than the base tier). Several agency owners in the communities we monitor describe reporting, not research, as the thing that justifies the price, and our experience matches: client-ready reports on a schedule remove a real block of monthly labor. The backlink audit with its built-in disavow workflow rounds this out. Ahrefs will show you the links and can export a disavow file, but deliberately leaves the toxicity judgment to you; Semrush scores them and walks you through acting on the score. If you do penalty-recovery work that workflow is genuinely useful, though most sites without a manual action never need to disavow anything. One caveat the community repeats, and [our own comparison](https://geotoolbox.ai/blog/semrush-vs-ahrefs) confirms: SEOs who run Semrush often still pay for Ahrefs specifically for link data. The index refreshes slower, and noisy lost-link alerts are the recurring complaint. If backlink data is your primary job, this is not your primary tool.
![Semrush Site Audit dashboard for geotoolbox.ai: Site Health 87 percent, beta AI Search Health 88 percent, and a Blocked from AI Search panel.](/blog/semrush-review/semrush-site-audit-ai-search-health.jpg)
Site Audit in 2026: an AI Search Health score sits beside classic site health, with a dedicated panel for pages blocked from AI crawlers.
The rest is a mixed bag. Position Tracking is solid (daily, down to zip code level, now including AI surfaces we cover below). The content tools, SEO Writing Assistant included, are useful on Guru and up. The social media toolkit is an afterthought, and nothing we would pay separately for. The learning curve is real, and it belongs next to the price. The interface has dozens of sections, and community threads about switching tools consistently describe weeks, not days, to feel fluent. Budget onboarding time like you budget the subscription. ## Accuracy: What Our Three Sites Show Reviews disagree wildly about Semrush's data accuracy. [One popular review](https://bloggingx.com/how-accurate-is-semrush/) measured its own site and found Semrush underestimating traffic by 43 percent. Another says the accuracy "resembles Google Search Console closely." Each ran a single site, which is exactly why they disagree. We compared Semrush's organic traffic estimates with Google Search Console clicks and Ahrefs estimates on three sites of very different sizes. Same week, same pull, checked again twelve days later, and the re-pull is a finding on its own: the models move. Over those twelve days Semrush's estimated traffic for our small site went from about 1,000 to 1,700 while Ahrefs revised its own estimate from about 1,500 down to 375, swings [we documented pull by pull](https://geotoolbox.ai/blog/semrush-vs-ahrefs) in the comparison piece. The table below uses the later pull for everything.
SiteSemrush est. monthly organicAhrefs est. monthly organicGSC clicks (30 days)
Small site (our own, a few months old)~1,700~375382
Mid-size site (manufacturer)~5,100~2,7003,703
Large site (marketplace)~397,300~316,00083,499
**TESTED, August 2026:** tool figures are each dashboard's headline estimate for the site's primary market: worldwide for our half-French small site, the US database for the two US-centric client sites, and worldwide throughout for Ahrefs. GSC counts Google web clicks over the trailing 30 days, unfiltered. Full methodology in [the comparison piece](https://geotoolbox.ai/blog/semrush-vs-ahrefs). Three patterns worth your attention. The error is large, and it is never the same size twice. Semrush read high on all three of our sites, but by anywhere from more than a third over to nearly five times over, so no fixed discount you apply to the dashboard number will correct it. Ahrefs happened to land within seven clicks of the Search Console number on our small site, a coincidence you should not extrapolate: it undershot the mid-size site by about a quarter and ran nearly four times over on the marketplace. Published tests point the other way too: one reviewer measured a 43 percent underestimate. A reviewer testing one site can conclude almost anything, in good faith. Estimates and clicks measure different things, but the gap is still the point. Tool estimates model visits from ranking positions and clickstream panels; Search Console counts actual Google clicks. Some divergence is structural. What matters for a buyer is that the site-level number on the dashboard is a modeled figure that can be off by hundreds of percent, and other reviewers have caught it missing low by nearly half. Nothing in the interface warns you either way. Keyword counts are definitions, not measurements. Semrush credited our small site with about 1,500 ranking keywords worldwide on the August 1 pull (1.6K a day later, as the screenshot above shows); Ahrefs counted 42 in its default view. Neither is lying. They count different positions, databases, and thresholds. Comparing your keyword count in one tool against a competitor's count in another is meaningless. Our working rule after running both tools side by side: use Semrush for relative comparisons (you against competitors, measured in the same index) and Search Console for the truth about your own Google clicks. That caveat is about precision, not coverage. Semrush sees more of the query space than anything else we run, which is exactly why the relative comparisons are worth having. The [7-day free trial](/go/semrush-seo?ref=semrush-review-accuracy) is enough time to run this same Search Console comparison on your own site before any invoice lands.
![Semrush Top Pages report for geotoolbox.ai with per-page columns for traffic, keywords, LLM Prompts, and referring domains, showing 107 organic pages and 5 cited pages in the US database.](/blog/semrush-review/semrush-top-pages-llm-prompts.jpg)
Top Pages now carries an LLM Prompts column per URL, blending AI-search data into a classic organic report. This August 2 pull reads 107 organic pages against the comparison piece's 106 a day earlier, and its 5 cited pages are the US-database slice of the 25 the AI Visibility report counts sitewide.
## The AI Visibility Toolkit, Reviewed by People Who Build One Semrush's 2026 repositioning is built around one idea: a unified AI visibility and SEO platform. The app's left rail now has a full AI Visibility section, and the Domain Overview carries an AI Search panel next to the classic organic data. We build [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) tooling at geotoolbox ourselves. That makes this section a review of a competing product, as disclosed above, and it also means we can judge it as practitioners. The documented feature set is substantial. Per [Semrush's own documentation](https://www.semrush.com/kb/1626), the AI Visibility Toolkit tracks your brand across **ChatGPT, Google AI Overviews, Google AI Mode, Gemini, and Perplexity**. It produces a [Visibility Score](https://www.semrush.com/kb/1493) (an index of your brand's mentions against the median of your top competitors), Share of Voice against named rivals, Prompt Research to find the questions people ask AI engines, and cited-source breakdowns showing which pages earn the citations. Daily Prompt Tracking, 25 custom prompts on the standalone toolkit's base tier, lets you benchmark competitors and watch shifts over time. On our own account we opened the Visibility Overview, Prompt Research, and the citation reports; we did not stress-test the Visibility Score or Share of Voice methodology. One clarification to the review ecosystem: several popular reviews claim Perplexity is not covered. Semrush's documentation says otherwise, and the confusion is understandable, because the Distribution by LLM panel inside our dashboard currently lists only ChatGPT, AI Overviews, AI Mode, and Gemini. We could not confirm Perplexity data on our own account, whether because our domain has no Perplexity citations yet or because the panel lags the docs. What is actually absent from the self-serve tiers: Claude, Grok, MS Copilot, and DeepSeek. Semrush's [Enterprise pricing page](https://www.semrush.com/pricing/enterprise/) sets it out in a comparison table: the AI Visibility Toolkit gets "ChatGPT, Google AI Overviews, Google AI Mode, Perplexity, and Gemini" with 25 to 200 prompts, while Enterprise's AI Optimization adds MS Copilot, Grok, Claude, and Deepseek on top, with unlimited prompts. **TESTED, August 2026:** the toolkit logged 40 AI citations across 25 cited pages for our site on August 2, up from 39 the day before (the dashboard's +2.6 percent badge) and from 31 citations across 19 cited pages in July. Directionally that matches the growth we see in our own logs. The absolute numbers do not: Bing Webmaster Tools reported 15.3K AI citations for the same domain over three months. The figures are not comparable: different periods, different surfaces, different definitions of a citation, and neither validates the other. Also worth reporting: on our domain the headline Visibility Score returned n/a and brand Mentions sat at zero across three of the four engines in the panel; the fourth, AI Mode, returned n/a, which is not the same claim as a zero. Only the citations counter produced data. On a domain our size the flagship score has nothing to chew on yet, which is itself a finding about who the toolkit is for. That is the right frame for this entire category, ours included: **AI visibility numbers are modeled samples, not ground truth.** Our own tracker samples a fixed prompt set on a fixed cadence and misses everything outside it; Semrush's samples differently and misses differently. AI answers are volatile and personalized, no vendor sees every response, and [our data study on AI search](https://geotoolbox.ai/blog/state-of-ai-search-2026) found the engines disagree with each other about who deserves a citation in the first place. Treat the scores as competitive direction, never as a count of anything. A tool that shows you trending up against named competitors on tracked prompts is useful. A tool that claims to measure your total AI presence is overpromising, whoever sells it. Where does that leave the toolkit? If you already live in Semrush, it is the most convenient way to get systematic AI tracking next to your SEO data. Pair it with first-party evidence (your logs, Bing Webmaster Tools, and a crawler-access check like our free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness)) and you have a defensible AI visibility practice.
![Semrush AI Visibility Overview for geotoolbox.ai showing 40 citations up 2.6 percent, 25 cited pages, zero brand mentions, and a Distribution by LLM panel listing ChatGPT, AI Overview, AI Mode, and Gemini.](/blog/semrush-review/semrush-ai-visibility-overview.jpg)
The AI Visibility overview on our domain: 40 citations across 25 cited pages, zero brand mentions, and four engines in the Distribution panel.
## Semrush Pricing in 2026: What You'll Actually Pay Semrush pricing got more complicated in 2026. There are now two front doors: the classic SEO plans, and the Semrush One bundles that fold the AI Visibility Toolkit into one subscription. All figures below are from [Semrush's pricing pages](https://www.semrush.com/pricing/seo-ai-search/) and its own [Semrush One documentation](https://www.semrush.com/kb/1624), re-checked on August 7, 2026.
PlanMonthly billingAnnual billing (per month)Limits
Pro (classic) — now sold as "SEO"$139$117.335 projects, 500 tracked keywords
Guru (classic) — off public pricing, see note$249.95$208.3315 projects, 1,500 keywords, content tools
Business (classic) — off public pricing, see note$499.95$416.6640 projects, 5,000 keywords, API
One Starter$199$165.175 sites, 500 keywords, 50 AI prompts daily
One Pro+$299$248.1715 sites, 1,500 keywords, 100 AI prompts
One Advanced$549$455.6740 sites, 5,000 keywords, 200 AI prompts
Each Semrush One tier inherits the features of its classic sibling (Starter maps to Pro, Pro+ to Guru, Advanced to Business) and adds the AI Visibility Toolkit on top. Every publicly sold plan starts with the same [7-day free trial](/go/semrush-seo?ref=semrush-review-pricing). One restructuring note, and it changed between our first check and this update: the public pricing page, now titled "SEO & AI Search Plans and Pricing," shows exactly four plans — a single "SEO" plan next to the three One bundles. That "SEO" plan is the classic Pro renamed, and very slightly repriced: $139 flat, down from Pro's $139.95. Guru and Business no longer appear on any public pricing page we could reach: they remain documented in Semrush's knowledge base and active for existing subscribers, but a new buyer today is being steered to the entry SEO plan or the bundles. If your budget was built around Guru, price the equivalent One tier instead. One more scope note before the catches: the classic plans are not entirely AI-blind. Semrush's own documentation shows they include sample-level AI citation and mention data, plus the AI Search Health score in Site Audit. What they cannot do is track prompts, which is the thing the toolkit actually sells. Now the parts the pricing page does not put in the headline. **The $99 AI toolkit gains nothing from paying annually.** The AI Visibility Toolkit standalone is $99 per month per domain, and [the toolkit page](https://www.semrush.com/pricing/ai/) opens on its Annually tab, where the fine print reads "billed annually." Switch the toggle to Monthly and the price is identical, $99, with that line gone and a different checkout offer behind the button. So unlike the SEO plans, where paying for a year takes $139 down to $117.33, committing twelve months to the toolkit buys you no discount whatsoever. We checked this at a signed-in checkout, not just on the pricing page: the monthly offer bills every 1 month, the annual one every 12, and the per-month figure is identical. What the toolkit genuinely lacks is a way to try it: per Semrush's own knowledge base there is no free trial of it at all, so the first $99 is a purchase, not a test. **Every plan includes one seat.** Additional users start at $45 per month each and scale up by plan. The $45 is a floor, not a rate: Semrush prices extra users at $45 on SEO and Starter, $80 on Pro+, and $100 on Advanced. A five-person team on Pro+ is not paying $299 a month, it is paying $299 plus four seats at $80, or $619 before a single add-on. **The add-on stack is where budgets die.** Local SEO, Traffic & Market Analytics, the Content Toolkit, Advertising, Social, and AI PR are all separately priced add-ons, several of them billed per user. The community's word for this is nickel-and-diming, and it's the most consistent pricing complaint we found. Price the tools you expect to open weekly before you commit. **The upgrade math favors One for a single-user account.** On monthly billing, Guru plus the $99 toolkit is about $349 against One Pro+'s $299, and both routes advertise a month-to-month option. Run it all-annual and the shape is the same: $307.33 against $248.17. [Semrush's own documentation](https://www.semrush.com/kb/1624) puts it plainly: an annual One Pro+ costs essentially the same as a monthly Guru (accepting the 12-month commitment annual billing implies), with quadruple the standalone toolkit's prompt allowance (100 on Pro+ versus 25) included. With Guru now pulled from the public pricing page, the bundle is the default route for new buyers anyway. Teams should still price their exact seat and domain count before switching, because per-user costs stack differently across the two routes. What is a prompt, in practice? One tracked prompt is one question checked daily across the covered AI engines. Fifty prompts on One Starter sounds like a lot until you split it: ten brand and product queries plus eight competitors at five prompts each exhausts it exactly. Agencies should budget prompts like they budget tracked keywords: they run out, and [Semrush's docs](https://www.semrush.com/kb/1493) price the refill at $60 a month for another 50, which is the add-on pattern again. The [free trial](/go/semrush-seo?ref=semrush-review-trial) runs seven days as standard per Semrush's own site (partner links sometimes extend it to 14; the length shown at your checkout is the one to calendar), and it requires a credit card. Given how Semrush handles cancellation, that matters more than usual. ## The Complaints Worth Taking Seriously Semrush holds a 4.4 rating from 4,024 reviews on [G2](https://www.g2.com/products/semrush/reviews) and a 1.7 from 1,354 on [Trustpilot](https://www.trustpilot.com/review/semrush.com), as rendered in Google's review snippets the day we checked. That split is not a contradiction. G2 reviews mostly discuss the product; Trustpilot's recent reviews mostly describe billing disputes. Both are telling you something true. **Cancellation is the big one.** Submitting the cancellation form is not the end of the process. [Semrush's own help documentation](https://www.semrush.com/kb/252-cancelling-your-account) states that after the form you receive a confirmation email whose link is valid for 24 hours, and "you must click this link to finalize your request." Miss that window and the subscription keeps renewing. Communities treat missed confirmations as a recurring trap. In the late-2025 through mid-2026 threads we monitored across r/SEO and adjacent subreddits, cancellation and renewal complaints were the most common Semrush topic, and [one long-running worth-it thread](https://www.reddit.com/r/SEO/comments/jc6llk/sem_rush_is_it_worth_it/) has proven durable enough to rank on page one for "semrush review" (US results, August 2026). The [Better Business Bureau profile](https://www.bbb.org/us/ma/boston/profile/marketing-software/semrush-inc-0021-553725) shows a D- rating, which BBB ties to 53 unanswered complaints, and the recent complaints on the profile center on trial charges and refused refunds. The refund fine print explains some of that heat: [Semrush's refund policy](https://www.semrush.com/refund) is a one-time, 7-day money-back guarantee that applies only to annual-or-longer commitments. Month-to-month subscriptions are not refundable at all, so the people angriest in those threads are often outside a window that never applied to them. If you [start the trial](/go/semrush-seo?ref=semrush-review-complaints): calendar the trial end date, cancel a few days early if you are not converting, and complete the confirmation email the same day it arrives. None of this is hard, but nobody at Semrush will do it for you. Support has thinned, by community account. Reports of template responses and slow escalation recur across 2025 and 2026 threads, usually attached to billing cases rather than product questions. Our own support contacts have been fine for product issues, so we mark this one community-reported rather than tested. **Then there is Adobe.** Adobe [completed its acquisition of Semrush](https://news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition) on April 28, 2026, an all-cash deal announced at $1.9 billion, and Semrush now sits in Adobe's Customer Experience Orchestration business. Practitioner sentiment in the threads we monitored skews strongly negative, with three recurring fears: enterprise-oriented pricing, slower product velocity, and the fate of self-service accounts inside a company built on suite deals. What has changed so far, from inside a paying account: nothing we can observe. Headline prices on the plans we hold did not move at the close (though Semrush restructured its packaging during 2026, as covered above), the product kept shipping (the AI Visibility rebrand landed post-announcement), and Adobe publicly committed to continued investment. The fears are about trajectory, not the present. We wouldn't cancel over the logo, but we also wouldn't sign a multi-year commitment on the assumption that pre-acquisition pricing survives the integration. In the threads we read, the complaints clustered on billing rather than the product itself, and data-quality complaints were rare (though our own accuracy testing above suggests the estimates deserve more scrutiny than the community gives them). People complain about paying for Semrush, rarely about using it. ## Who Should Buy It (And Who Shouldn't) ### Agencies and client-facing teams Yes. The combination of breadth, PPC intelligence, and white-label reporting is unmatched among the tools we have run, and per-client project workflows are what the platform is built around. Budget for seats honestly and the invoice still beats stitching four tools together. [Start the trial](/go/semrush-seo?ref=semrush-review-agencies) on your busiest client account. ### In-house teams at mid-size companies Yes, with the seat math done first. If two or more people live in SEO data weekly and you touch paid search at all, Semrush earns its slot. If your whole SEO program is one person maintaining rankings, a cheaper tier of a simpler tool covers most of what you need. ### Solo operators and small sites Usually no. This is where the community's "not worth it yet" consensus is right. The [free plan](/go/semrush?ref=semrush-review-free) covers ten reports a day, enough to check an idea, and it costs nothing to keep around while you grow into the paid tiers. Past that, tools like [SE Ranking](/go/seranking?ref=semrush-review-alts) or [Mangools](/go/mangools?ref=semrush-review-alts) cover keyword research and tracking at a fraction of the price (our [Semrush alternatives](https://geotoolbox.ai/blog/semrush-alternatives) guide picks the best by job), and [free SEO tools](https://geotoolbox.ai/blog/best-free-seo-tools) cover audits and basics. Come back to Semrush when clients or scale arrive. ### AI-visibility-first buyers Check the bundle before the standalone. If AI search tracking is your primary goal, [One Starter at $199](/go/semrush-one?ref=semrush-review-aifirst) is the cheapest full Semrush package with AI tracking included, and the [standalone $99 toolkit](/go/semrush-ai?ref=semrush-review-aifirst) only wins if you need nothing from the SEO side. It advertises a monthly option, but there is no trial and no annual discount, so the first $99 is a purchase. We compared the wider field in our [AI visibility tools roundup](https://geotoolbox.ai/blog/best-ai-visibility-tools) if Semrush is not a foregone conclusion. That field includes our own [geotoolbox](https://geotoolbox.ai/pricing) — from $99 a month with a 7-day trial, reaching eight engines on higher tiers, including Claude, Copilot, and Grok — three engines the self-serve toolkit does not cover — reviewed in that roundup with the same disclosure made there. ## Is Semrush Worth It in 2026? For agencies and teams that will use its breadth: yes, clearly. It is the strongest all-in-one we have paid for, and nothing else we have used combines organic, paid, and AI search data this deeply in one subscription. For everyone else it is a utilization question, not a quality question: count the toolkits you would open in a normal week. At three or more, buy it; at one, do not; at two, let the trial decide. The thing that would change this verdict is Adobe. If the next renewal cycle arrives with enterprise pricing and a thinner self-service tier, the math above stops holding, and we will update this review to say so. If you are ready to test it, [start the 7-day Semrush free trial](/go/semrush-seo?ref=semrush-review-closing) and run it on your own data, or [compare the Semrush One bundles](/go/semrush-one?ref=semrush-review-closing) if AI visibility is part of the plan. Run it against your Search Console the way we did. Separately, if you want the AI-search side checked without a subscription, our free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) takes about a minute. Whichever way you decide, set that cancellation reminder first. ## Frequently Asked Questions ### What is Semrush? Semrush is an all-in-one digital marketing platform covering keyword research, competitor analysis, site audits, rank tracking, PPC research, and, since 2025, AI search visibility tracking. It has been owned by Adobe since April 2026 and sits alongside Ahrefs as one of the two dominant paid SEO suites. ### Is Semrush worth the money? For agencies and teams using three or more of its toolkits, yes. The breadth, PPC data, and reporting justify the price. For solo operators with one small site, usually not: cheaper tools cover the basics, and Semrush's per-seat and per-add-on pricing punishes light use. The [7-day free trial](/go/semrush-seo?ref=semrush-review-faq) is the cheapest way to settle it on your own data. ### How much does Semrush cost per month? The entry SEO plan (the classic Pro, now sold as "SEO") runs $139 monthly. The Semrush One bundles, which add the AI Visibility Toolkit, cost $199, $299, and $549. Annual billing cuts up to 17 percent; the entry plan drops to $117.33 a month. Guru ($249.95) and Business ($499.95) still exist for current subscribers but no longer appear on the public pricing page. Extra user seats start at $45 per month, and most specialty toolkits are separately priced add-ons. ### How do I cancel a Semrush subscription? Submit the cancellation form from your subscription settings, then click the link in the confirmation email Semrush sends; per Semrush's help docs the link is valid for 24 hours and the cancellation is not final until you click it. Cancel a few days before renewal and complete the confirmation the same day. ### How accurate is Semrush's traffic data? Treat it as directional. Across our three test sites, Semrush overestimated every time, from more than a third over Google Search Console clicks to nearly five times over, while published single-site tests have measured meaningful underestimates too. Use it to compare competitors within the same index, and use Search Console for the truth about your own Google clicks. ### Does Semrush track AI search visibility? Yes. Semrush documents AI Visibility Toolkit coverage of ChatGPT, Google AI Overviews, AI Mode, Gemini, and Perplexity, with a Visibility Score, Share of Voice, Prompt Research, and daily prompt tracking; on our own account the dashboard showed data for the first four. It costs $99 per month per domain, the same figure on either billing term and with no free trial, or comes bundled in Semrush One plans. ### Is Semrush better than Ahrefs? For breadth, yes; for backlink data, no. Semrush covers more surface (PPC intelligence, content tools, AI visibility) and sees more of the query space, while Ahrefs refreshes link data faster and ran closer to Search Console in our three-site accuracy test. Many teams run both; if you must pick one, choose by your primary job. The full head-to-head data is in our [Semrush vs Ahrefs comparison](https://geotoolbox.ai/blog/semrush-vs-ahrefs). ### Is there a free version of Semrush? Yes, with real limits: one demo project and ten report requests per day. It works for checking a domain or a keyword occasionally. The free trial of the paid plans requires a credit card, so set a reminder before it converts. ## Sources - Semrush pricing page - Semrush, checked August 2026 - `semrush.com/pricing/` - AI Visibility Toolkit pricing - Semrush, checked August 2026 - `semrush.com/pricing/ai/` - What is Semrush One? - Semrush Knowledge Base, checked August 2026 - `semrush.com/kb/1624` - AI Visibility platform coverage - Semrush Knowledge Base, checked August 2026 - `semrush.com/kb/1626` - Prompt Tracking limits - Semrush Knowledge Base, checked August 2026 - `semrush.com/kb/1493` - Cancelling your account - Semrush Knowledge Base, checked August 2026 - `semrush.com/kb/252-cancelling-your-account` - Cancellation and refund policy - Semrush, checked August 2026 - `semrush.com/refund` - SEO & AI Search plans and pricing (replaced the SEO Classic plans page) - Semrush, checked August 7, 2026 - `semrush.com/pricing/seo-ai-search/` - Enterprise plans (AI Visibility toolkit vs AI Optimization LLM-coverage table) - Semrush, checked August 7, 2026 - `semrush.com/pricing/enterprise/` - How Accurate Is Semrush (Detailed Tests & Analysis) - Akshay Hallur, BloggingX - `bloggingx.com/how-accurate-is-semrush/` - Adobe to acquire Semrush (deal terms) - Adobe Newsroom, November 19, 2025 - `news.adobe.com/news/2025/11/adobe-to-acquire-semrush` - Adobe completes Semrush acquisition - Adobe Newsroom, April 28, 2026 - `news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition` - Semrush reviews - Trustpilot, checked August 2026 - `trustpilot.com/review/semrush.com` - Semrush reviews - G2, checked August 2026 - `g2.com/products/semrush/reviews` - Semrush, Inc. profile - Better Business Bureau, checked August 2026 - `bbb.org/us/ma/boston/profile/marketing-software/semrush-inc-0021-553725` - SEM Rush: Is it worth it? - r/SEO, Reddit - `reddit.com/r/SEO/comments/jc6llk/sem_rush_is_it_worth_it/` --- ## DeepSeek V4: What It Actually Is Now (GA, Pricing, and the Catches) > DeepSeek V4 explained, current to August 2026: what changed at GA, Flash vs Pro, the real API cost, self-hosting, and how honest the benchmarks really are. - Canonical: https://geotoolbox.ai/blog/deepseek-v4 - Published: 2026-08-01 · Updated: 2026-08-23 Ask the major AI assistants what DeepSeek V4 is and you get a different answer from each, most of them wrong. With web search switched off, ChatGPT calls it "rumors and roadmap speculation," Gemini dates itself to 2025 and says no such model exists, and Claude declines to answer. Turn search on and it barely improves: ChatGPT lands on the April preview announcement and concludes V4 has not left the preview stage, which stopped being true weeks ago. In one line: DeepSeek V4 is Chinese lab DeepSeek's current open-weight model family, a fast Flash model and a flagship Pro, generally available since August 2026 under an MIT license. Here is the fuller version, checked against primary sources as of August 2026: what DeepSeek V4 actually is now that it has shipped, what it genuinely costs once you account for the parts the sticker price hides, whether the benchmark claims survive an independent look, and whether you can put it anywhere near real work. Where a number is DeepSeek's own claim rather than an outside measurement, we say so, because on a model this new that distinction is most of the story. ## What DeepSeek V4 Is (and What Changed at GA) **DeepSeek V4 is the current family of open-weight models from [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek), the Chinese lab in Hangzhou, and its flagship V4-Pro reached general availability on August 13, 2026** after a months-long preview that opened on April 24. If you read an explainer that still calls V4 a preview, it is out of date, and so is DeepSeek's own V4-Pro model card, which still carries preview wording (more on that under the benchmarks). V4 is not a single model. **DeepSeek-V4-Pro** is the flagship, built for hard reasoning, agentic work, and demanding code. **DeepSeek-V4-Flash** is the fast, cheap workhorse for everything else. Both are [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) models, both ship a one-million-token context window, and both are released as [open weights](https://geotoolbox.ai/blog/open-weights-vs-open-source) under the MIT License, so you can download and run them yourself. Several things changed when V4 moved from preview to GA, and each one trips people up. It retired the old `deepseek-chat` and `deepseek-reasoner` model names on July 24. It quietly refreshed Flash to a new build, DeepSeek-V4-Flash-0731, on July 31, so the model you call today is not quite the one reviewers tested in May. And on August 16, 2026 it made good on months of warnings about a "significant" price increase, replacing the flat all-day rate with the peak/off-peak split covered below. One limit to set expectations early: V4 is text in, text out. There is no image, audio, or video input on the mainline models, which is the clearest single gap against ChatGPT, Gemini, and Claude (DeepSeek shipped an experimental V4-Flash-Vision-Exp on August 21, 2026, but it is a test model, not the GA release). If your work needs to read a screenshot or a chart, V4 is not the tool. ## The V4 Lineup: Flash vs Pro Flash and Pro are not a scaled-down copy and its bigger sibling; they are different mixture-of-experts configurations at very different price points. The rule of thumb about which to use also shifted after GA, once the refreshed Flash got strong enough to handle work people used to send to Pro.
ModelSize (total / active)Best forAPI price (per 1M tokens)
V4-Flash284B / 13B activeChat, extraction, classification, summarization, high-volume agent loops, most coding$0.22 in / $0.66 out off-peak, up to $0.44 / $1.32 at peak
V4-Pro1.6T / 49B activeHard reasoning, complex multi-step analysis, the most demanding code$0.66 in / $1.98 out off-peak, up to $1.32 / $3.96 at peak
Both are mixture-of-experts designs, which is why the active parameter count is a fraction of the total: the model holds a huge number of parameters but switches on only a small slice for any given [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai). That is the trick behind the low running cost. The default is **Flash for almost everything.** It is cheap enough that at normal volumes the bill is a rounding error, and on DeepSeek's own July 31 numbers the refreshed Flash matches or beats the earlier Pro preview on agent and coding tests. **Reserve Pro** for the work that genuinely needs it, hard reasoning and the most demanding code, where the quality gain is worth roughly triple the output cost. If you cannot tell whether a job needs Pro, it probably does not; start on Flash and move up only when you hit a wall. The July 31 refresh matters here. That build, DeepSeek-V4-Flash-0731, is what the wave of "Flash just got better" reviews are reacting to; a high-reasoning variant of it placed seventh overall, and third among open-weight models, on the Frontend Code Arena leaderboard at the end of July. The calling name (`deepseek-v4-flash`) did not change, so nothing breaks, but the model behind it did. If you benchmarked Flash before August, benchmark it again. ## The 1M Context Window Is Real, but It Is Lossy The headline feature is the one-million-token context window, and the interesting part is how DeepSeek made it affordable rather than the number itself. V4 uses a hybrid attention design, combining what DeepSeek calls Compressed Sparse Attention and Heavily Compressed Attention. At the full million-token setting, the technical report puts V4-Pro at roughly 27% of the per-token inference compute and about 10% of the KV-cache memory of the previous generation, DeepSeek-V3.2. That is a real engineering win, and it is why V4 can quote a 1M window at a price a dense model could not reach. There is a second change agent builders care about: in tool-using conversations, V4 preserves its reasoning across user turns instead of dropping it when a new message arrives, so multi-step workflows hold together better. One caller obligation comes with it: you have to pass that earlier reasoning back on each request, or the API rejects the call. Here is the catch the spec sheet leaves out. **A window that is cheap to run behaves as if it is lossy, and you feel it as the model "ignoring" things.** Developers running V4 on large inputs report, consistently, that it seems to forget instructions placed near the top of a long prompt, and some say the drop-off starts well before the million-token ceiling. DeepSeek has not published a per-position retrieval breakdown, and the technical report does not call the compression lossy, but the most plausible explanation is the same attention compression that makes the window affordable. Either way, the fix is a workflow habit, not a bug report. Put your load-bearing instructions near the end of the prompt, not buried at the top, and repeat any constraint that absolutely must hold. Treat the 1M window as a large working memory that is strong at "find the relevant passage" retrieval and weaker at holding a rule you stated tens of thousands of tokens earlier. Used that way, the context window is a genuine advantage; trusted blindly, it is where V4 disappoints people. ## How Good Is It, Really? DeepSeek's own numbers are strong. Its V4-Pro model card, the same one that still labels V4 a preview, reports 80.6% on SWE-bench Verified, 93.5% on LiveCodeBench, and 90.1% on GPQA Diamond, and it frames its knowledge results as an open-source record. Those are worth knowing. They are also vendor-reported, on a high-effort "Max" setting, and they are exactly the numbers you should check against someone independent. Independent evaluations tell a more grounded story. [Artificial Analysis](https://artificialanalysis.ai/models/deepseek-v4-pro), which runs its own suite, places V4-Pro at an Intelligence Index of 44, ranking it sixth among the models it tracks, above the median for comparable models but not at the frontier. The US government's AI standards body, CAISI, went further, re-running DeepSeek's headline benchmarks and adding held-out ones of its own.
Benchmark (CAISI test)DeepSeek V4-ProGPT-5.5
SWE-Bench Verified74%81%
PortBench (held-out coding)44%78%
CTF-Archive (cybersecurity)32%71%
GPQA-Diamond90%96%
A couple of things jump out. First, CAISI measured V4-Pro at 74% on SWE-bench Verified, against the 80.6% DeepSeek publishes, and stated plainly that "DeepSeek V4 scores better on DeepSeek's self-reported evaluations than on CAISI evaluations." That is the vendor-benchmark gap in one sentence. Second, CAISI's overall read is that V4's capabilities lag the frontier by about eight months, with the widest gaps on held-out coding and cybersecurity, where training-data contamination cannot inflate the score. A couple of caveats keep this fair. CAISI's math results run the other way, with V4-Pro at 97% on OTIS-AIME-2025, near parity with the US models. And CAISI ran its tests back in April, on the preview build, not the GA model; DeepSeek's own preview-to-GA numbers show sizable gains from re-training alone, so part of the gap is likely build drift rather than pure inflation, and nobody has yet published an independent re-test of the shipped model.
![Grouped bar chart from NIST CAISI's independent test showing DeepSeek V4-Pro trailing GPT-5.5 on four benchmarks: SWE-Bench Verified 74 vs 81, PortBench 44 vs 78, CTF-Archive 32 vs 71, and GPQA-Diamond 90 vs 96.](/blog/deepseek-v4/caisi-benchmark-gap.png)
DeepSeek V4-Pro vs GPT-5.5 on NIST CAISI's independent benchmarks. The gap is widest on held-out coding and cybersecurity, and V4-Pro's measured SWE-bench score (74) sits below DeepSeek's own 80.6 claim.
None of that makes V4 bad. On cost-per-task, CAISI compared it against GPT-5.4 mini, the nearest cheap US model, and found V4 ran cheaper on five of seven benchmarks but anywhere from 53% less to 41% more expensive depending on the task, a direct illustration of the thinking-token tax covered below. An independent scored test by coding-tool maker Kilo put V4-Pro at 77 out of 100, between Claude Opus 4.7 at 91 and Kimi K2.6 at 68, though that test predates the current Claude and Kimi flagships. The honest summary: **V4 is a frontier-class model on price and a near-frontier model on capability, strongest on math, reasoning, and long-context retrieval, and clearly behind the closed leaders on the hardest real-world coding, cybersecurity, multimodal work, and the longest agentic loops.** If a page tells you it simply beats GPT-5.5 and Claude, it is selling the vendor's numbers, not the measured ones. ## What DeepSeek V4 Really Costs The web and mobile apps are free. The money question starts on the API, and this is where V4's reputation as "the cheap one" is earned. Here are the current per-million-token rates. For a fuller cost breakdown, see our [DeepSeek pricing](https://geotoolbox.ai/blog/deepseek-pricing) guide.
ModelInput (cache miss)Input (cache hit)Output
V4-Flash — off-peak$0.22$0.007$0.66
V4-Flash — peak$0.44$0.014$1.32
V4-Pro — off-peak$0.66$0.022$1.98
V4-Pro — peak$1.32$0.044$3.96
A few catches separate that sticker price from your actual bill. **The price increase DeepSeek was signaling has landed.** On August 16, 2026, DeepSeek replaced its flat all-day rate with the peak/off-peak split shown above: peak hours (01:00-04:00 and 06:00-10:00 UTC, Monday to Friday, with weekends billed at off-peak) bill at exactly double the off-peak rate across every column. In a market where these prices mostly fall, that is a rare reversal, and it is worth building your budget around the peak share of your actual traffic rather than assuming off-peak applies to everything. **Thinking mode is on by default, and it bills at the output rate.** Before it answers, V4 generates internal reasoning tokens, and those count as output even though you never see them. A short reply can quietly produce thousands of billed tokens first. If your bill is higher than the sticker suggested, this is almost always why. Turn thinking off for routine work (classification, extraction, short replies) and leave it on only where the reasoning earns its cost. **Cache hits are where the real economics live.** When the start of a request matches something processed recently (a reused system prompt, a document you keep sending), those tokens bill at the cache-hit rate, dropping Flash input from $0.22 to $0.007 off-peak (or $0.44 to $0.014 at peak). Caching is automatic and free; you do not mark anything as cacheable, you just keep the stable part of your prompt at the front. Structure prompts to reuse a common prefix, and the same workload can cost a fraction of the headline number. There is also no separate long-context surcharge, so a million-token prompt bills at the same per-token rate as a short one at whatever time it runs. ## How to Access V4: App, API, or Self-Host There are a few ways in, and the cheapest is not always the obvious one. The **free web and mobile app** at chat.deepseek.com gives you the full V4 chat with web search, no payment. The **direct API** is the cheapest paid route and the one to build on. **OpenRouter** and other resellers are convenient, but they route to third-party hosts whose rates differ from DeepSeek's own, so a price you see there may not match DeepSeek's. If cost is the priority, go direct, though DeepSeek's own API can be capacity-constrained under load, which is part of why the resellers exist. If you build on the API, migration is the thing that will bite you. The old model names retired on July 24, 2026 at 15:59 UTC, and code that still calls them now fails.
Old name (retired)Call this insteadWhat it maps to
deepseek-chatdeepseek-v4-flashFlash, non-thinking mode
deepseek-reasonerdeepseek-v4-flashFlash, thinking mode
The drop-in swap has a catch. Many people assume `deepseek-reasoner` was the Pro-tier reasoning model and that swapping it keeps that power. It did not, and it does not: `deepseek-reasoner` mapped to Flash in thinking mode, not to Pro. If you want Pro-level reasoning you have to ask for `deepseek-v4-pro` explicitly, or you are silently running a smaller model. **Self-hosting is the real escape hatch, with a hardware asterisk.** Because the weights are MIT-licensed, you can run V4 on your own machines, which is the cleanest answer to any data-residency concern. In practice, V4-Pro at 1.6 trillion parameters is out of reach for individuals and hard even for well-funded teams; it wants a multi-GPU server with hundreds of gigabytes of memory even when quantized. V4-Flash at 284 billion is the realistic self-host target, the 0731 build a touch larger once its speculative-decoding module is counted, and it is what most of the [run-it-locally](https://geotoolbox.ai/blog/run-llm-locally) discussion is about. Plan for a high-memory multi-GPU node served with vLLM or SGLang rather than the llama.cpp or LM Studio hobbyist path. Local tooling was rough at launch; quantized builds and framework support improved through the summer, so check the current state of your runtime before you commit. If you have run [Kimi K3 locally](https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally), the constraints will feel familiar. ## Can You Use It at Work? This is the question that decides most professional adoption, and the answer is genuinely "it depends," not a yes or a no. If you use DeepSeek's hosted API or app, your prompts go to servers in China. For many enterprises, government contractors, and regulated industries, that alone is a hard stop, and plenty restrict it outright. That is a routing and jurisdiction decision, separate from anything about the model's quality. There is a middle route the open weights make possible: buy V4 from a Western host. Microsoft's Azure AI Foundry, Fireworks, and Together all serve V4-Flash and V4-Pro from US infrastructure, so your prompts never touch DeepSeek's servers and you get an enterprise contract and a data-residency commitment without owning a GPU cluster. You pay a margin over DeepSeek's direct rates; that is the price of jurisdiction. Or self-host. Because V4 is MIT-licensed, a team that cannot send data anywhere it does not control can run the model on its own infrastructure, so prompts never leave the building and the hosted app's logging and terms do not apply. One caveat self-hosting does not fix: political refusals and topic-alignment are trained into the weights, so a self-hosted V4 still dodges or skews on sensitive questions the way the hosted one does. Self-hosting answers the data-residency question, not the will-it-answer-straight one. The wider trust picture, including how DeepSeek handles data on its own servers, is covered in our [DeepSeek explainer](https://geotoolbox.ai/blog/what-is-deepseek). ## DeepSeek V4 vs the Competition The framing that keeps people straight is that V4 is a model you can download, while ChatGPT, Gemini, and [Claude](https://geotoolbox.ai/blog/claude-vs-chatgpt) are products that rent you access to a closed model. On raw token price, V4 sits in a different bracket, undercutting the Western flagships by roughly an order of magnitude, though the gap narrows on finished tasks once thinking tokens are counted. Against the current closed leaders from OpenAI, Google, and Anthropic (including [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5)), the trade is the familiar one: cost and openness versus polish, multimodal range, reliability across varied tasks, and a lead on the hardest agentic coding. Against the other open-weight Chinese models, V4 is the reasoning-and-price specialist, [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3) leans more agentic, and [Qwen](https://geotoolbox.ai/blog/what-is-qwen) trades on breadth and multilingual coverage. For a side-by-side, our guide to the [Chinese models compared](https://geotoolbox.ai/blog/chinese-ai-models-compared) puts them in one table, and the [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking places V4 in the full field. ## Quirks and Gotchas Worth Knowing A few things the benchmarks never mention still surprise people. **V4 does not reliably know it is V4.** Ask it which model it is and it may answer V3; its self-report is not dependable. Do not use the model's own answer to check which version you are running; check the API model name you called instead. **Creative and open-ended work regressed slightly.** Users who write fiction or do open-ended ideation report a stronger positivity bias and harder-to-steer tone than V3.2 had. It is a narrow, named-audience complaint rather than a headline flaw, but if creative writing is your use case, test before you switch. **Agent instruction-following is contested.** Some builders report V4 acting outside an explicit scope in agent harnesses; others report none. The consensus fix is architectural: give it a clear orchestrator and tight tool definitions rather than trusting a long freeform instruction. **R2 still has not shipped.** The reasoning model people keep asking about, DeepSeek-R2, was widely expected in 2025 on the back of press reports and has not landed; V4's built-in thinking mode is DeepSeek's answer for now. Treat R2 as a rumor until a model card exists. ## What This Means for Getting Cited by AI Come back to the opening. The reason those AI assistants gave conflicting, mostly wrong answers about V4 is not that V4 is obscure; it is a widely covered launch. It is that the assistants are only as current as the sources they can reach, and most of them had not caught up. That is the whole game for anyone who wants to be represented accurately by AI. When a model, a product, or a fact is new, the engines fall back on whatever stale material they can find. In practice, getting quoted correctly comes down to the unglamorous disciplines this page tries to model: date your claims, cite each number to its source, put the direct answer next to the heading, and re-check the volatile facts on a schedule. That is the [foundation of generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization). Stale answers come from stale or unreachable sources, so the first move is making sure the engines can even reach and parse your pages. Run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see where the AI crawlers get blocked or trip up on your site before the next launch makes the question urgent again. ## Frequently Asked Questions ### Is DeepSeek V4 out yet, or is it a preview? It is out. V4-Pro reached general availability on August 13, 2026, ending a preview that ran from April 24. Explainers and AI answers that still call it a preview are out of date, and so is DeepSeek's own V4-Pro model card, which still carries preview wording. ### Should I use V4 Flash or V4 Pro? Start with Flash for almost everything: chat, extraction, summarization, high-volume agent loops, and most coding. Move to Pro only for hard reasoning, complex multi-step analysis, or demanding code, where the quality gain is worth roughly triple the output cost. ### How much does DeepSeek V4 cost per month? The app is free. On the API your bill tracks your token volume, so the honest answer is to work it out from the table above: multiply your monthly input and output tokens by the Flash rates, and remember thinking mode adds billed reasoning tokens on top. Most individual use lands in a few dollars a month and scales from there; Pro runs about triple Flash for the same volume. Our [DeepSeek pricing](https://geotoolbox.ai/blog/deepseek-pricing) guide has worked examples. Cache hits and turning thinking off for simple tasks move the bill most. ### Is DeepSeek V4 about to get more expensive? It already has. On August 16, 2026, DeepSeek replaced its flat all-day rate with a peak/off-peak split: off-peak, V4-Flash runs $0.22 input / $0.66 output per million tokens and V4-Pro $0.66 / $1.98, and during peak hours (01:00-04:00 and 06:00-10:00 UTC) every rate exactly doubles. That was the "significant increase" the pricing page had been signaling since late July. Budget on the off-peak rates and add the peak premium for whatever share of your traffic runs in those seven hours a day; DeepSeek's pricing has moved twice this year, so re-check the [official page](https://api-docs.deepseek.com/quick_start/pricing) before committing a budget. ### Why is my DeepSeek bill higher than the sticker price? Almost always because thinking mode is on by default and its reasoning tokens bill at the output rate, even though you never see them. Turn thinking off for routine tasks, and structure prompts to reuse a common prefix so more tokens hit the cheaper cache-hit rate. ### I used deepseek-chat or deepseek-reasoner. What do I switch to? Both retired on July 24, 2026. Call `deepseek-v4-flash` for the old chat and reasoner behavior. Note that `deepseek-reasoner` mapped to Flash in thinking mode, not to Pro, so if you want Pro-level reasoning you must call `deepseek-v4-pro` explicitly. ### What is the cheapest way to use DeepSeek V4? The direct DeepSeek API, not a reseller. OpenRouter and similar marketplaces add their own margin on top of DeepSeek's rates. Going direct plus earning cache hits is the lowest-cost path. ### Is the 1M-token context window real? Yes, but it is lossy. The compression that makes a million tokens affordable means detail from early in a long prompt gets summarized, so the model can seem to ignore instructions placed at the top. Put critical instructions near the end and repeat anything that must hold. ### Can I use DeepSeek V4 at work? Through the hosted API or app, your data goes to servers in China, which many companies and regulated industries prohibit. The MIT-licensed open weights let you self-host instead, keeping data in-house, which is the usual path for sensitive work. ### Is DeepSeek V4 open source? The weights are open under the MIT License, so it is open-weight and free to download, run, and fine-tune. That is not the same as fully open source, since the training data and pipeline are not released. ### What happened to DeepSeek R2? It has not shipped. R2 was expected in 2025 and remains unreleased; V4's built-in thinking mode is DeepSeek's current reasoning offering. Treat any R2 claim as a rumor until a model card exists. ## Sources - DeepSeek API Models & Pricing (rates, model names, forthcoming price-increase notice) - DeepSeek, accessed August 2026 - `api-docs.deepseek.com/quick_start/pricing` - DeepSeek-V4-Pro model card (architecture, parameters, vendor benchmarks, license) - DeepSeek on Hugging Face - `huggingface.co/deepseek-ai/DeepSeek-V4-Pro` - DeepSeek-V4 technical report (hybrid attention, long-context efficiency) - arXiv - `arxiv.org/abs/2606.19348` - CAISI Evaluation of DeepSeek V4 Pro (independent benchmarks, frontier-lag finding) - NIST, May 2026 - `nist.gov/news-events/news/2026/05/caisi-evaluation-deepseek-v4-pro` - DeepSeek V4 Pro - Intelligence, Performance & Price Analysis (independent index) - Artificial Analysis - `artificialanalysis.ai/models/deepseek-v4-pro` - We Tested DeepSeek V4 Pro and Flash (independent scored review) - Kilo - `blog.kilo.ai/p/we-tested-deepseek-v4-pro-and-flash` - DeepSeek V4 Flash and V4 Pro in Microsoft Foundry (Western hosting, US/EU residency) - Microsoft - `techcommunity.microsoft.com/blog/azure-ai-foundry-blog/introducing-deepseek-v4-flash-and-v4-pro-in-microsoft-foundry/4515174` - Three reasons why DeepSeek's new model matters - MIT Technology Review, April 2026 - `technologyreview.com/2026/04/24/1136422/why-deepseeks-v4-matters` - DeepSeek-V4-Pro community discussion (long-context and creative-writing feedback) - Hugging Face - `huggingface.co/deepseek-ai/DeepSeek-V4-Pro/discussions/179` --- ## Best AI SEO Agencies (2026): An Honest Buyer's Guide > The best AI SEO (GEO) agencies in 2026, sorted by job, plus how to verify any agency's citation claims yourself in 20 minutes and when not to hire one at all. - Canonical: https://geotoolbox.ai/blog/best-ai-seo-agencies - Published: 2026-07-30 · Updated: 2026-08-11 More of your buyers now start their research with ChatGPT or Perplexity than with Google, and if you are not in those answers, you are not on the shortlist. Demand for [AI SEO](https://geotoolbox.ai/blog/ai-seo) agencies has grown to match. This guide sorts the credible ones by the job they are best at, then shows you something most agency-written lists will not: how to pressure-test any firm's citation claims yourself, and when hiring one is the wrong move. One disclosure up front. We are geotoolbox, an AI visibility tracker with a small GEO service, so we are on this list too, in last place, with a real drawback spelled out. Fair warning while you read: this guide often lands on "measure first, and a tool may be enough," and tracking is what we sell, so weigh the advice against how self-limiting it turns out to be. The rest is the whole picture with the hype removed. ## What an AI SEO Agency Actually Does An AI SEO agency helps your brand show up when people ask ChatGPT, Perplexity, Google's AI Overviews, or Gemini for a recommendation. The job title changes depending on who you ask. Some call themselves a GEO agency (generative engine optimization), some an [answer engine optimization](https://geotoolbox.ai/glossary/answer-engine-optimization) agency, some an LLM SEO agency. Across the AI engines we tested for this guide (ChatGPT, Perplexity, Gemini, and Google's AI Overviews, queried in July 2026), those terms describe the same job. The demand behind this is real, not hype. In [G2's 2026 research](https://learn.g2.com/g2-2026-ai-search-insight-report), 51% of B2B software buyers now start their research with an AI assistant rather than Google, up from 29% a year earlier, and one in three ended up buying from a vendor they had not previously heard of. Yet [CommonMind found](https://www.commonmind.com/blog/state-of-ai-visibility-in-b2b-saas) that while 93% of B2B SaaS marketers call AI visibility critically important, only 14% report a mature, documented strategy for it. That gap between urgency and readiness is the market these agencies serve. The work itself has shifted. Traditional SEO chases a ranking. AI SEO chases a mention: being the source an AI names, cites, or recommends inside its answer. That is a different target, and it needs a different set of deliverables. A real AI SEO engagement usually spans the following: - **An AI visibility audit and ongoing monitoring** across each engine, not a single snapshot - **Content restructured for extraction**, so an answer-first passage can be lifted cleanly into a generated response - **Entity and schema work**, so models can connect your brand to the right topics and facts - **Topical authority**, the same depth-and-coverage play SEO has always rewarded - **Digital PR and third-party citations**, because models weight what other sites say about you - **Prompt-level tracking**, measuring a frozen set of buyer questions on a schedule One point every engine agreed on: GEO complements traditional SEO, it does not replace it. AI systems still pull from the web index, so crawlability, structured data, and content quality remain the foundation. If you want the terms straight, our explainer on [GEO vs AEO vs SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) untangles them. The rest of this guide covers who does this work well, what it costs, and how to check that any agency's claims are real before you sign. ## Is "AI SEO" Just Rebranded SEO? An Honest Answer If you have spent five minutes in a marketing forum, you have seen the accusation: AI SEO is old SEO with a find-and-replace on the sales deck. It is worth taking seriously, because it is partly true. The fundamentals have not moved. [Google's own guidance](https://developers.google.com/search/docs/appearance/ai-features) is blunt that there is no secret trick for AI features. The same helpful content, structured data, and crawlable pages that won rankings are what make you eligible to be cited. An agency that sells you a brand-new "AI methodology" while ignoring your technical foundation is selling you a label. But the work on top of that foundation is genuinely new, and the data shows why. In Ahrefs' [study of 863K keywords](https://www.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637/), the share of AI Overview citations coming from top-10-ranked pages fell from 76% in July 2025 to 38% eight months later. Ranking in the top 10 no longer guarantees you get cited. Being cited now depends on how a passage reads to a model, whether your entity is clear, and whether other sources corroborate you, none of which a rank-tracking dashboard measures. It is worth keeping the hype in check from the other direction too. In 6sense's 2025 Buyer Experience Report, nearly all buyers used AI models during the journey, but mostly to summarize and organize research, not to replace it, and they still averaged around sixteen interactions with the vendor they chose. AI is reshaping how buyers discover and shortlist, not vaporizing the rest of the funnel. An agency selling you a world where AI does everything is overselling as badly as one pretending nothing changed. So the short answer: the base is SEO, and the new layer is real. A legitimate agency respects both. A repackaged one charges for the rename. One piece of the new layer that most teams miss entirely is reachability, whether AI crawlers can even fetch your pages, which we cover in our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers). ## How We Ranked These Agencies (and Why There Is No Single "#1") One note on method, since this category is full of lists that hide theirs. We are on this list ourselves (disclosed above), in last place, with a real drawback, and not first. We deliberately do not assign a composite score out of 50 or 100. When we ran the same buyer prompts across ChatGPT, Perplexity, Gemini, and Google's AI Overviews in July 2026, each engine cited a materially different set of agencies. A single blended number would paper over that variance and imply a precision that does not exist. It is also the exact move that makes most "best agency" scorecards untrustworthy: the author sets the criteria, then wins its own contest. That matters because the category is noisy. When we looked at the wider pool of content ranking and being cited for this query, most of it was agency-owned blog posts and a scattering of low-authority content farms, with few genuinely independent sources. The incentive to publish a self-flattering roundup is enormous, and buyers have caught on, which is why many now assume these lists are seeded until proven otherwise. We would rather earn the opposite assumption. One signal did hold up. Across those July 2026 runs, iPullRank was the one agency that showed up consistently across more than one engine. Everyone else moved around depending on which model we asked, which is itself the lesson: there is no universal leaderboard, only a best fit for your segment and budget. Instead of a score, we weighed several things: whether a firm actually surfaced across the engines when we ran the prompts; whether its claimed strengths held up across the other agency roundups we cross-checked, not just its own marketing; how it fares against the pain points real buyers report in communities; and which segment it genuinely fits. Where a firm is strong in one lane and weak in another, we say so in the "watch for" instead of averaging it into a misleading total. So we have sorted the list by job, given each firm a plain verdict, and flagged what to watch. Treat it as a shortlist to pressure-test, not a leaderboard. ## The Best AI SEO Agencies in 2026, by Use Case
AgencyBest forPricing modelWatch for
iPullRankEnterprise technical GEORetainer, ~$12k+/moEnterprise-only floor
OnelyEnterprise, full-journey GEORetainer, ~$10k+/moNot built for SMB budgets
Omniscient DigitalB2B SaaS content-led growthRetainerContent-first, lighter on technical
First Page SageThought-leadership GEORetainerFintech-leaning; check vertical fit
Siege MediaContent and digital PR that earns citationsRetainer, from ~$11k/moLight on technical GEO
Seer InteractiveData and analytics-driven SEOHourly / project, $10k+ minIntegrated shop, not an AI specialist
TinuitiEnterprise full-funnel performanceCustom / enterpriseAI sits inside a broad program
Single GrainGrowth-stage, multi-channelRetainer, $10k+ minGeneralist breadth over GEO depth
SearchbloomMid-market, flexible termsRetainerRanks itself #1 on its own list
OptimistB2B SaaS specialistsRoadmap from $7.5k; retainer from $3k/moNarrow ICP ($10M-$500M ARR)
Coalition TechnologiesEcommerce and broad-marketRetainerSelf-claims "#1"; generalist
geotoolboxMeasurement-first, verify before you buy$925 audit; retainer from $3,500/moSmall, founder-led; not enterprise full-service
A note on the price figures below: they are floors as published by each firm or reported in the roundups we cite, not quotes. Treat them as starting points and confirm on your own call. ### iPullRank - best for enterprise technical GEO The one name that held up across more than one AI engine in our tests. [iPullRank](https://ipullrank.com/) built its reputation on deep technical SEO and now frames AI work as "relevance engineering," the retrieval-layer craft of making content extractable and entities unambiguous. If you run a large, complex site and your blocker is architecture rather than blog volume, this is the strongest fit. Watch for the price: this is an enterprise engagement, typically starting around $12,000 a month, and the AI-specific practice is newer than the technical SEO it rests on. ### Onely - best for enterprise, full-journey GEO [Onely](https://www.onely.com/) (which surfaced in Google's AI Overview for this query when we checked) pairs technical execution with mapping the full customer journey rather than isolated keywords. Good for enterprise brands that want one partner across discovery to conversion. The same caveat as iPullRank applies: pricing starts around $10,000 a month, so it is not built for a small-business budget. ### Omniscient Digital - best for B2B SaaS content-led growth If your growth model is content and your buyers are other businesses, [Omniscient](https://beomniscient.com/) has the case studies, including large organic gains for SaaS brands. Their strength is editorial depth and topical authority, which travels well into [AI citations](https://geotoolbox.ai/glossary/ai-citation). Watch for the flip side: it leans content-first, so pair it with technical help if your foundation needs work. ### First Page Sage - best for thought-leadership GEO [First Page Sage](https://firstpagesage.com/) leans into authority and thought leadership as the engine for citations, and it shows up in AI answers for exactly that positioning. Strong if you sell on expertise, particularly in finance-adjacent categories. Watch for vertical fit: the playbook is tuned for B2B authority plays and is a weaker match for transactional or local businesses. ### Siege Media - best for content and digital PR that earns citations [Siege](https://www.siegemedia.com/) is a content and link machine, and in an AI world those earned mentions on third-party sites are exactly what models weigh. If your gap is off-site authority and citations, this is a proven pick. Watch for the balance: Siege leans content and PR over the technical and entity side of GEO, so it works best when your site fundamentals are already solid. Engagements commonly start around $11,000 a month. ### Seer Interactive - best for data and analytics-driven SEO [Seer](https://www.seerinteractive.com/) is an analytics-first agency that folds GEO into a broader, measurement-heavy practice. A good fit for data-mature teams that want AI search inside a rigorous reporting culture. Watch for specialization: Seer is a full-service shop, not an AI-only boutique, so if you need a dedicated GEO team, ask exactly who does that work. ### Tinuiti - best for enterprise full-funnel performance [Tinuiti](https://tinuiti.com/) is one of the largest independent performance agencies, best known for managing media and organic together at scale for retail and consumer brands. AI search sits inside that full-funnel offering rather than standing alone. The fit is enterprise brands that want AI visibility managed alongside paid, by a team that already handles their broader marketing. Watch for scale: a boutique GEO problem can get lost inside a program this broad, so make sure the AI scope is named explicitly in the contract and someone owns it. ### Single Grain - best for growth-stage, multi-channel teams [Single Grain](https://www.singlegrain.com/) wraps SEO and content into a wider growth program, pairing organic with paid and content strategy, which suits growth-stage companies that want several channels under one roof and one point of contact. Watch for depth: breadth is the selling point, so if GEO is your single most important problem, a dedicated specialist will likely go deeper than a generalist juggling four channels. ### Searchbloom - best for mid-market with flexible terms [Searchbloom](https://www.searchbloom.com/) is a well-reviewed mid-market SEO shop with a documented framework and flexible engagement terms, which lowers the risk of a long lock-in. Watch for the obvious, which to its credit Searchbloom flags itself: it publishes its own ["best AI SEO companies" list](https://www.searchbloom.com/strategy/best-ai-seo-agency-companies-services-usa/) that ranks Searchbloom first, and says so openly on the page. The work looks solid; read that self-ranking the same way you should read this one. ### Optimist - best for B2B SaaS specialists [Optimist](https://www.yesoptimist.com/) is narrowly focused on B2B SaaS and is unusually transparent about pricing, publishing a roadmap fee and month-to-month retainers rather than hiding everything behind a call. It reports strong LLM-referral results for software clients. Watch for the ICP: the model is tuned for SaaS companies in a specific revenue band, so it is a weaker fit outside that lane. ### Coalition Technologies - best for ecommerce and broad-market Coalition runs a productized AI SEO and GEO offering with a large case-study library and strong ecommerce experience. Watch for the marketing: [Coalition describes itself](https://coalitiontechnologies.com/ai-seo-company), in its own page metadata, as "the #1 AI SEO (GEO) agency," which is a self-made claim, not a measurement or a third-party award. Treat it as a capable generalist and verify the specifics for your niche. ### geotoolbox - best for measurement-first buyers who want to verify before they buy Full disclosure, this is us. geotoolbox is an AI visibility tracker with a small, founder-led done-for-you service. Our angle is measurement: we start with a fixed-price [AI visibility audit](https://geotoolbox.ai/services/ai-seo-agency) at $925 — below the $1,000-to-$10,000 one-off-audit band later in this guide — so you can see where you actually stand across engines before committing to anything ongoing, and our free tools let you run the baseline yourself. The retainer, from $3,500 a month, is one of the two lowest published figures on this list, because the tracker does the measurement work a larger team bills hours for. Watch for a couple of real limits. We are the newest and smallest name here, with the thinnest public case-study record. And we sell a tracker, so we have a standing incentive to point you toward ongoing measurement, which is exactly why we tell you to run the baseline yourself first. We also left ourselves out of the cross-engine test above, since scoring ourselves on our own list would prove nothing, and we are only one of many small, founder-led GEO shops that could sit in this lane. If you need an enterprise full-service team or a big content-production engine, several firms above are a better fit, and we will say so. ## What AI SEO Agencies Cost in 2026 Few agencies publish rates, so buyers get quoted anything from under $1,000 to $50,000 a month with no anchor. Here is the range we saw hold up across several agency pricing guides and our own engine research.
EngagementTypical costWhat it buysTime to results
One-off audit$1,000-$10,000 (one-time)Baseline across engines, gap analysis, roadmapDays
Small / local$1,000-$5,000Technical fixes, schema, a few optimized pagesCitations in 2-4 weeks
Mid-market$5,000-$12,000Full stack: content, entity, digital PR, per-engine trackingMovement in 3-6 months
Enterprise$12,000-$50,000+Technical depth at scale, multi-team programsRevenue impact 6-12 months
Freelancer$75-$150/hr or $1,500-$4,000/moOne specialist, narrower scopeVaries with capacity
In-house hire$70,000-$110,000/yr + toolsFull control, dedicated focusSlow to hire, then ongoing
A couple of numbers are worth committing to memory. If you already pay for SEO, bolting GEO onto that retainer commonly runs a [20-30% uplift](https://www.demandlocal.com/blog/how-to-price-geo-services/) rather than a second full invoice, though scoped add-ons run higher, so ask what a bigger delta actually buys. And on the low end, [Searchbloom](https://www.searchbloom.com/strategy/best-ai-seo-agency-companies-services-usa/) treats anything under roughly **$1,000 a month** as a warning sign, usually thin, automated work rather than real strategy. If you are wondering why two agencies quote wildly different numbers for the same brief, it comes down to a few cost drivers: how many pages you need optimized, whether new content creation is included or just optimization of existing pages, how much off-site citation and digital PR work is in scope, and how deep the entity and technical work goes. A quote is really a scope, so compare what each price actually buys, line by line, not the headline figure. Two "$5,000 a month" proposals can describe completely different amounts of work. The timelines matter as much as the price. Technical fixes can produce citation pickup within a few weeks, but meaningful movement takes 3 to 6 months, and revenue impact 6 to 12. Any agency promising full results much faster than that is overpromising. A sensible way to de-risk the spend is to start with a one-off audit rather than a twelve-month retainer, so you buy a baseline before you buy a commitment. There is also an ongoing-cost reality that pricing pages gloss over: AI citations tend to be less durable than a hard-won ranking. Models favor fresh sources, so a page that earns citations can quietly slide out of answers over the following months if nothing keeps it current. Part of any retainer therefore goes to defending citations you already have, not just earning new ones. That is normal, but it is worth asking an agency to split those out, so you know how much of your spend is maintenance rather than growth. ## How to Verify an Agency's Results Yourself (The 20-Minute Test) Here is the part few agency-written listicles will teach you, because it lets you check their work. You do not need to take a citation screenshot on faith. You can reproduce any agency's core claim yourself, before or during an engagement, and the quick version takes about twenty minutes. The trap is that AI answers are non-deterministic. Ask ChatGPT the same question twice and you can get different brands. That means a single screenshot proves almost nothing, and it is exactly how a weak agency cherry-picks a lucky session. The fix is to test for consistency, not for one good result.
![Four-step process to verify an AI SEO agency's claims: freeze your buyer prompts (5 for a quick check, 20 to 30 for a full baseline), run each several times in fresh logged-out sessions across ChatGPT, Perplexity, Gemini and AI Overviews, score how consistently you appear per engine, and separate mentions, citations, and recommendations.](/blog/best-ai-seo-agencies/20-minute-verification-test.png)
Verify any agency's claims yourself: score consistency across engines, not a single screenshot.
There is a fast version and a thorough one. **The 20-minute triage.** Pick 5 buyer-intent prompts, the real questions a customer asks ("best [your category] for [use case]"). Run each in 3 fresh, logged-out sessions across the pair your buyers actually use, usually ChatGPT and Perplexity. That is thirty queries in about twenty minutes, enough to catch an agency showing you one lucky screenshot. **The full baseline** takes about half a day, or a tracker runs it for you on a schedule. Freeze 20 to 30 prompts and run each across all of them, 3 to 5 rounds apiece: ChatGPT, Perplexity, Gemini, and Google's AI Overviews. Logged-out matters, or your own history skews the answer. Either way, score the same things: 1. **How often you actually appear**, not whether you appeared once. As a rough read, showing up in under a fifth of runs means you are effectively invisible, a fifth to a half is fragile, and consistent presence is the goal. Movement only counts if you measure the same prompts before and after. 2. **Per engine, never blended.** A brand can hold real [share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) in Perplexity and be invisible in Gemini the same week, because [each engine cites a different source mix](https://geotoolbox.ai/blog/state-of-ai-search-2026). One combined number hides that spread. 3. **Mention, citation, or recommendation.** Being named in passing is a mention. Named with a source link is a citation. Named as the answer is a recommendation, and only the last reliably moves revenue. An agency that reports "visibility went up" without saying which one moved is not telling you enough. Whatever an agency reports should survive you re-running these prompts. If it does not, that tells you what you needed to know. You can run the triage by hand or start with our free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) and go from there. Either way, the number is yours to reproduce, not theirs to assert. ## Proving It Actually Drove Pipeline Visibility is not the finish line, pipeline is, and AI search makes that link hard to see. Most AI referrals are effectively zero-click: someone reads a recommendation in ChatGPT, then shows up days later by typing your name into Google. So the measurement stack has to be layered, and an agency that cannot describe it is guessing. Four signals, used together, get you close enough to brief a CFO: - **Segment AI referral traffic in GA4.** Filter for `chatgpt.com`, `perplexity.ai`, Gemini, and Copilot referrers so AI-sourced sessions stop hiding inside direct. - **Add a "how did you hear about us" field** to your forms. Self-reported attribution catches the zero-click journeys that no tracking script can follow. - **Watch branded search and direct traffic for lift** in the weeks after your citations climb. A rise with no new campaign is often AI discovery cashing out. - **Tie your AI-bot server logs to accounts that convert.** Standard analytics miss most AI bot hits because those bots do not run JavaScript, so reachability lives in the logs, not the dashboard. No single signal is clean. Together they let you say whether the spend is working, which is more than "our visibility score went up" will ever tell you. ## Red Flags: Spotting a Repackaged or Overpriced Agency The engines we tested converged on the same warning signs, and they line up with what burned buyers say in marketing forums. Any one of these should slow you down. - **"Guaranteed rankings" or "guaranteed AI citations."** No one controls a probabilistic model's output. A guarantee is a sales tactic, not a capability, and it was the most-cited red flag in our tests. - **They cannot name the platforms they monitor.** If an agency says it "optimizes for AI search" but cannot tell you it tracks ChatGPT, Perplexity, Gemini, and AI Overviews separately, there is no real measurement underneath. - **Reporting that is only rankings and traffic.** In a world where [less than a third of searches send a click](https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/), a dashboard with no AI-specific metric is measuring the wrong thing. - **Suspiciously cheap retainers.** "AI SEO" under about $1,000 a month is usually automated output with no human review, the "AI slop" buyers complain about after the fact. - **A second full retainer on top of your existing SEO.** GEO layered onto SEO commonly runs a 20-30% uplift, sometimes more. Being charged for a whole second service for work that overlaps your current one is the double-billing pattern buyers describe as predatory. - **A single-screenshot case study.** As the 20-minute test shows, one lucky citation is not evidence. Ask for consistency across runs. - **A "best AI SEO agency" list the agency wrote to include itself.** Whether it ranks first or, like this one, discloses its own entry lower down, the incentive is the same. Buyers have started flagging these as seeded content designed to teach AI models a fake consensus. Read every self-including list, this one included, with that in mind. - **An agency that offers to build that for you.** Some will sell you seeded roundups and fake-consensus content as a "citation" service. It can work briefly, but it is a brand risk that communities and platforms are actively policing, not a durable strategy. The common thread is a firm that sells a story it will not let you check. ## Questions to Ask Before You Sign Bring these to the first call. The good answers are specific; the bad ones are vague or defensive. 1. **Show me a live prompt where your own agency, or a named client, appears in ChatGPT or Perplexity right now.** Not a screenshot from March. A firm that does this work can reproduce it in the meeting. 2. **What is in this scope that is not already in my existing SEO retainer?** This is the anti-double-billing question. A straight answer separates a real add-on from a rename. 3. **Do you report per engine, or as one blended score?** The right answer is per engine, because the source mixes differ. A single number is a warning. 4. **Do you track mentions and citations separately?** Being named and being linked are different outcomes. An agency that conflates them is not measuring carefully. 5. **Give me a first-90-days plan, and a named condition under which you would tell me it is not working.** A willingness to define failure is the strongest signal of an honest partner. 6. **How do you track which AI bots actually crawl my pages?** Standard analytics miss most AI bot traffic because those bots do not run JavaScript, so this needs server or edge-level logging. If they have not thought about it, they are guessing at reachability. 7. **How much of the monthly retainer defends existing citations versus earning new ones?** AI citations decay as content ages, so maintenance is real work. A firm that has an answer has been doing this long enough to have seen the decay. 8. **What are the terms: minimum commitment, notice period, and who owns the prompt set, tracking data, and content if we part ways?** This is where lock-in hides. You want to leave with your data and your pages, not start over. This is the diligence the category now demands. ## When You Should Not Hire an AI SEO Agency Few agency-authored lists include this section, because every quote is a yes. Here are the cases where the answer is not yet, or not at all. **Your buyers are on Google Maps, not ChatGPT.** If you run a local business and your demand is "coffee shop near me," people are still tapping the map pack, not asking an assistant for a shortlist. Spend on local SEO first. AEO can wait until your category actually gets asked about in AI. **Your technical foundation is not fixed.** Paying for off-site citation building while AI crawlers cannot even reach your pages is buying the roof before the walls. Fix crawlability, structured data, and answer-first content first. That is often in-house work, and it is the layer everything else depends on. **Your category has no AI demand yet.** Some niches barely surface in generative answers. Run the 20-minute test on your own buyer prompts. If no engine is returning brands like yours, there is little for an agency to optimize toward right now. **You can run the baseline and basics yourself.** If your team can freeze a prompt set, fix schema, and write clear answer-first pages, you may not need a retainer at all, just a tracker to watch progress. We wrote a full breakdown of that trade-off in [GEO services versus software](https://geotoolbox.ai/blog/geo-services-vs-software). **You need everything, not just GEO.** If organic, paid, social, and PR are all gaps, a GEO-only boutique is the wrong shape. Hire a broader team, or a specialist for the one channel that is actually holding you back. Knowing when to walk away is worth more than any shortlist. An agency that agrees you are not ready yet is one worth remembering for when you are. ## Agency, Freelancer, In-House, or a Tool An agency is one of several ways to get this work done, and it is not always the right one. **An agency** (typically $3,000 to $12,000 a month) buys breadth and a team. It fits when GEO is a real priority, your budget supports it, and you would rather manage outcomes than tasks. The risk is paying agency rates for work a single person could do. **A freelancer** ($75 to $150 an hour, or a light monthly retainer) is cheaper and often just as skilled for a focused scope. The trade-off is capacity and continuity: one person gets sick, gets busy, or moves on. **An in-house hire** ($70,000 to $110,000 a year plus tools) gives you control and full-time focus, which suits companies where AI visibility is a permanent, strategic channel. The cost is time: hiring a genuinely good GEO practitioner in a two-year-old field is slow. **A tool-first approach** is the cheapest and the most underrated. If your team can run the baseline, fix the basics, and act on what a tracker shows, you may not need any of the above yet, just measurement. This is the path we build for at geotoolbox, and we would rather tell you to start here than sell you a retainer you are not ready for. The [best AI visibility tools](https://geotoolbox.ai/blog/best-ai-visibility-tools) guide compares the trackers, and our [agency-versus-software breakdown](https://geotoolbox.ai/blog/geo-services-vs-software) walks through the decision in more depth. Most teams end up combining these over time: measure with a tool, hire a specialist or agency for the heavy lift, then keep watching with the tool. The order that rarely works is signing a long retainer before you have measured anything at all. ## The Short Version Choosing well in this category comes down to a few moves. Measure your baseline first, so you know where you actually stand across engines. Use the list above as a shortlist, matched to your segment and budget. Then pressure-test any pitch with the 20-minute triage and the questions above before you sign anything. If you want a rule of thumb: an enterprise site whose blocker is technical points to iPullRank or Onely; a B2B SaaS team that grows through content points to Omniscient; an ecommerce brand points to Coalition; a buyer who wants the measurement verified before committing points to our own $925 audit; and if your budget is under a few thousand a month, do not hire yet, run the baseline yourself first. That is a starting point, not a verdict, and anyone who promises you more certainty than that is overselling. If you want to start with the measuring, you can run the baseline yourself for free. Our [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) shows where you stand and whether AI crawlers can even reach you, which is the first thing most agencies should be fixing anyway. If you would rather have the work done for you, our [founder-led GEO service](https://geotoolbox.ai/services/ai-seo-agency) is one option among the good firms above, and we will tell you if you are better served by a tool or by someone else on this list. ## Frequently Asked Questions ### What is an AI SEO agency? An AI SEO agency helps your brand appear in AI-generated answers from ChatGPT, Perplexity, Google's AI Overviews, and Gemini. The terms GEO agency (generative engine optimization) and AEO agency (answer engine optimization) describe the same job. The work spans technical fixes, content restructured for extraction, entity and schema work, digital PR, and per-engine citation tracking. ### How much does an AI SEO agency cost in 2026? Retainers typically run from about $1,000 a month for small businesses to $12,000 or more for enterprise programs, with mid-market engagements around $5,000 to $12,000. One-off audits run $1,000 to $10,000. Adding GEO to an existing SEO retainer usually adds a 20-30% uplift, sometimes more. Anything under roughly $1,000 a month is generally automated work, not real strategy. ### Are AI SEO agencies worth it? It depends on your situation. An agency is worth it when AI visibility is a real priority, your buyers actually use AI to research your category, and your technical foundation is sound. It is not worth it if you are a local business whose buyers use Maps, if your site fundamentals are broken, or if your team can run the baseline and basics in-house. Measure first, then decide. ### Is GEO just rebranded SEO? Partly. The fundamentals, technical health, structured data, and content quality, are the same ones traditional SEO always rewarded, and Google's own guidance says there is no separate trick. But the new layer is real: ranking in the top 10 no longer guarantees a citation, so per-engine tracking, entity work, and answer-first content are genuinely different skills. A legitimate agency respects both; a repackaged one charges for the rename. ### How do I verify an agency's AI citation claims? Run the 20-minute triage: freeze 5 buyer-intent prompts (20 to 30 for a full baseline), run each several times in fresh, logged-out sessions across the engines your buyers use, and score how consistently your brand appears rather than trusting a single screenshot. Read the results per engine, and separate mentions, citations, and recommendations. Anything an agency reports should survive you reproducing it yourself. ### Can my current SEO agency do AI SEO, or do I need a separate one? Sometimes your current agency can, if it already does strong technical SEO and is genuinely tracking AI engines. Ask what is in the AI scope that is not already in your retainer, and whether it reports per engine. If the answer is vague or it wants a full second retainer for overlapping work, that is the double-billing pattern to avoid. A specialist is only worth it when it does something your current partner demonstrably cannot. ## Sources - G2 - The Answer Economy: 2026 AI Search Insight Report - `learn.g2.com/g2-2026-ai-search-insight-report` - CommonMind - The 2026 State of AI Visibility in B2B SaaS - `commonmind.com/blog/state-of-ai-visibility-in-b2b-saas` - 6sense - 2025 B2B Buyer Experience Report - `6sense.com/science-of-b2b/buyer-experience-report-2025` - Ahrefs AI Overview citation study, via Search Engine Journal - `searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637/` - SparkToro - Less Than a Third of Google Searches Send a Click (2026) - `sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/` - Google Search Central - AI features and your website - `developers.google.com/search/docs/appearance/ai-features` - Fuel Online - AI SEO Pricing 2026 - `fuelonline.com/ai-seo-geo/ai-seo-pricing-aeo-geo-cost-2026/` - Demand Local - How to Price GEO Services - `demandlocal.com/blog/how-to-price-geo-services/` - Searchbloom - Best AI SEO Companies (pricing + red-flag threshold) - `searchbloom.com/strategy/best-ai-seo-agency-companies-services-usa/` --- ## How to Run Kimi K3 Locally: What It Actually Takes (2026) > How to run Kimi K3 locally: the real hardware for a 1.56TB model, five routes ranked by cost and speed, and an honest verdict on self-hosting. - Canonical: https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally - Published: 2026-07-28 · Updated: 2026-08-14 The day after Moonshot released Kimi K3's weights, we asked three assistants how to run it. Queried on July 28, 2026 with web search off, Claude Opus 4.8 said it had no reliable information that a model called Kimi K3 exists and offered to help with K2 instead. Gemini 2.5 Flash got it backwards, insisting K3 is proprietary, cloud-only, and that "download size is irrelevant because you don't download the weights." ChatGPT with search on answered correctly. The weights are 1.56TB, sitting public on Hugging Face right now. So the real question about the largest open-weight model ever released is not whether you can have it, but how you run it. Here is what that takes, route by route, with the numbers. ## The Short Answer: Can You Run Kimi K3 Locally? Not on a normal machine. Kimi K3 is a [2.8-trillion-parameter](https://geotoolbox.ai/blog/what-is-kimi-k3) model, and its weights ship as a 1.56TB download. Loading them takes more fast memory than any workstation, gaming rig, or single Mac holds. Moonshot recommends production deployment on a supernode of 64 or more accelerators, eight 8-GPU nodes or more, though it has never published a minimum viable configuration. But local runs on a spectrum, from pointing your editor at a hosted endpoint to buying a rack. The question is which of these routes matches your actual constraint: privacy, cost, speed, or curiosity.
![Bar chart to scale showing fast memory each option holds against the 1.56TB Kimi K3 needs: an RTX 4090 (24GB) is a thin sliver, a 512GB Mac Studio and a 1TB-RAM server both fall short of the dashed 1.56TB line, and only an 8x B300 datacenter node (about 2.3TB) clears it.](/blog/how-to-run-kimi-k3-locally/kimi-k3-memory-gap.png)
The weights must sit in fast memory to run at full speed. A RAM-based rig can still run K3 slowly by offloading experts to system memory (Route 3), but a datacenter node is the only thing that clears the line outright.
RouteWhat you needRealistic speedRough costBest for
5. Use a providerNothing, an API key~37-148 tok/s, provider-dependent$3 in / $15 out per M tokensAlmost everyone
1. Rent a K3-capable node8x B300 or 16x B200, by the hourSame as Route 2~$40-$100+ / hourTrying it, short jobs
2. Self-host on vLLMOne 8x B300 node (or 16x B200)111 tok/s TP8, 331 with spec decodingDatacenter capex or cloudOrgs with real GPU infra
3. CPU + RAM offloadBig-RAM server + 1-4 GPUs~10 tok/s decode, slow prefill~$5k-$65k one-timeHomelab, overnight batch
4. Mac Studio cluster4-5x 512GB M3 Ultra~15 tok/s by analogy, unmeasured~$40k+ one-timeUnified-memory experiments
Speeds for the self-hosted rows are the ceiling on the right hardware, not what a first build hits. Most people should land on a provider (Route 5), or a rented node (Route 1) to test first. The rest of this guide is why, and what the others cost. ## Why It Takes So Much: The Memory Math The number that decides everything is 1.56TB, the size of the [Kimi K3 repository](https://huggingface.co/moonshotai/Kimi-K3) on Hugging Face, spread across 96 weight shards. A smaller figure, around 594GB, gets repeated in some launch coverage. The arithmetic rules it out: 594GB would imply about 1.2 trillion parameters at 4-bit, and K3 has 2.8 trillion. Whatever that number measured, it is not this model's weights. The naive calculation is 2.8 trillion parameters times 4 bits, which comes to 1.4TB. The actual repo is 1.56TB because MXFP4 stores a shared scale factor per block and the non-expert layers sit above 4-bit. So budget 1.56TB resident just for weights, and 1.8TB to 2TB once you add the KV cache and compute buffers. That is the wall, and it does not shrink. Because K3 is quantization-aware-trained at 4-bit (with 8-bit MXFP8 activations), there is no free halving from the usual 16-bit-to-4-bit step waiting for you; an 8-bit build is bigger than what shipped, not smaller. The [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) design is the second trap. K3 activates only 16 of its 896 routed experts (plus two shared experts) per token, so it computes like a much smaller model. That saves compute, not memory. The full weight set still has to be resident, because any token might route to any expert. A sparse model gives you a cheaper forward pass, never a smaller footprint. Then there is the [context window](https://geotoolbox.ai/glossary/context-window). At 1,048,576 tokens, a filled context adds real memory pressure on top of the weights. Kimi Delta Attention, the hybrid design K3 uses, keeps a constant-size state for most layers and softens that cost, but a long-context job still needs headroom you have to plan for. One more thing the self-host audience needs before building: K3 ships under a custom [Kimi K3 License](https://geotoolbox.ai/blog/open-weights-vs-open-source), which permits commercial use but is not OSI open source. A model-as-a-service operator whose revenue passes $20 million over any 12 months must negotiate a separate agreement with Moonshot, and any product above 100 million monthly users or $20 million in monthly revenue must display "Kimi K3" in its interface. For a team self-hosting to serve customers, that is a go/no-go detail, not trivia. ## Route 1: Rent a K3-Capable Node (Try Before You Buy) Before you spend a dollar on hardware, rent the model for an hour. This is the route the local-inference community reaches for first: at launch, before third-party providers came online, launch-week reports had Moonshot's own hosted K3 rate-limiting people mid-task, so renting was the reliable way to evaluate it before committing to a build. The catch is what you rent. An 8x H100 box at roughly $20 an hour is the default SKU people reach for, but 8x H100 is 640GB, which cannot load a 1.56TB model at any precision this guide endorses. You need the same floor as Route 2, an 8x B300 or 16x B200 node, which runs closer to $40-$100+ an hour on RunPod, Lambda, or Vast. You pay for the hour, not the tokens, which is exactly why this route is for measuring rather than production. Load your real workload, not a toy prompt. Push your actual context length and concurrency through the rented node, measure throughput and time to first token, then decide whether owning hardware makes sense. An hour of rental is cheaper than a wrong build. If you only want K3 in your editor, [a provider](#route-5-just-use-a-provider) is cheaper still. ## Route 2: Serve It Yourself With vLLM This is the route Moonshot designed K3 for, and the only one that reaches full speed. [vLLM](https://vllm.ai/blog/2026-07-27-k3) shipped day-0 support and has the fullest documentation, though SGLang is a supported alternative. The serving stack is the least of your problems. The hardware is. Moonshot's 64-accelerator supernode is a production-serving recommendation; the minimum config that merely fits the weights is different. Per the vLLM team, the floor is one 8x B300 node (or a GB300 NVL72), or a minimum of 16x B200 or GB200 on the previous generation, and AMD's MI355X carries ROCm support from launch if you are not on NVIDIA. Their own note is blunt: the model "can barely fit in a single NVIDIA DGX B300." No single H100, H200, or B200 holds it, so any real deployment is a distributed cluster, and the interconnect matters as much as the cards. Tensor parallelism over a slow fabric is a downgrade, not an upgrade: without high-bandwidth RDMA between nodes, the all-to-all traffic stalls and prompt processing collapses. On GB300 NVL72, vLLM measured 111 tokens per second per user at tensor-parallel size 8 (331 with DSpark speculative decoding), and 118 at TP16, rising to 370 with DSpark. The serve command matters, because K3 needs flags most launch guides leave out: ```bash vllm serve moonshotai/Kimi-K3 \ --tensor-parallel-size 8 \ --max-model-len 131072 \ --trust-remote-code \ --load-format fastsafetensors \ --enable-prefix-caching \ --enable-auto-tool-choice \ --tool-call-parser kimi_k3 \ --reasoning-parser kimi_k3 ``` A few of those are easy to miss. Set `--max-model-len` below the full 1M (131072 is a safe start; raise it once you have measured KV-cache headroom) or the server tries to allocate cache for the entire context and OOMs on first launch. `--load-format fastsafetensors` needs the `fastsafetensors` package installed separately. Prefix caching is off by default for K3 because the hybrid cache design is still stabilizing, so pass the flag for the prefill savings, but validate output on repeated prefixes before trusting it in production. The `kimi_k3` parsers are required for tool use and K3's always-on reasoning, and because the chat template is a Python program rather than a Jinja file, trusting remote code is not optional. One silent trap the flags do not cover: for multi-turn conversations you must pass the full `reasoning_content` back on each request, or quality degrades across the conversation with no error. The realistic steps: 1. Provision the cluster (8x B300 minimum, or 16x B200) with high-bandwidth RDMA between nodes. 2. Pull the official vLLM container image; the dependency chain pulls pre-release libraries like FlashInfer, so bare-metal installs are painful. 3. Download the weights from Moonshot's official repository only. Plan for it: 1.56TB across 96 shards is hours on a fast line and most of a day on a home connection, needs staging disk separate from where it runs, and is a prerequisite for the homelab route too. 4. Serve with the command above, starting well below the 1M context. 5. Load-test with your real concurrency before pointing production traffic at it. ## Route 3: CPU + RAM Offload (The Homelab Route) If you cannot put 1.56TB in VRAM, you can put most of it in system RAM. This is the homelab pattern: a big-memory server with roughly 1TB to 1.5TB of DDR RAM and one to four GPUs, keeping the routing and attention layers plus the KV cache on the GPU and offloading the routed experts to system RAM. The engine for this is llama.cpp with its expert-offload flags (`-cmoe` to keep experts on the CPU, `--no-mmap` for better throughput at the cost of a much longer load), or ktransformers, which practitioners report hitting 8-11 tok/s with on the prior K2.6 line using a dual-socket EPYC plus a few consumer GPUs. Community estimates put K3 decode for this pattern around 10 tokens per second, which sounds usable. It is not usable for interactive work, and the reason is prefill. Generating tokens at 10 per second is fine; processing your prompt is where these rigs collapse. At roughly 1 token per second of prompt processing, a 20,000-token context takes almost six hours before the model replies. Overnight batch jobs survive that. An agent loop that re-reads a large context on every step does not. The clearest data point comes from deltafin, an [open-source experiment](https://github.com/gavamedia/deltafin) that reads K3's experts from local disk on demand on an Apple M1 Max with 64GB of memory. The author's published per-token budget is the lesson: the repo's benchmark now sits around 3.4 seconds per token (0.29 tok/s), roughly 20x faster than its first published version on July 27, 2026. Most of that time goes to moving weights off the drive and unpacking them, not to the compute itself. This is a fast-moving experiment, so check the repo for the current number, and profile the I/O path before you buy a faster GPU. A couple more numbers set expectations. Community estimates put a 2-bit build north of 800GB, so with KV cache and buffers you want at least 1TB of RAM to attempt it. And the real ceiling is memory bandwidth, not core count or capacity: builders report identical-RAM rigs differing 3x on throughput. Choose the platform by memory channels rather than cores, favor an EPYC or Xeon with enough CCDs and L3 to saturate the controller and with AVX-512 support, and avoid dual-socket NUMA penalties where one socket will do. One non-obvious constraint: at 5-10kW a full rig is 40-100A at the meter, which practitioners raise alongside dedicated circuits, cooling, and home insurance as real line items before the build ever runs. If you want a gentler on-ramp to local models first, our [guide to running an LLM locally](https://geotoolbox.ai/blog/run-llm-locally) covers the tooling on hardware you already own. ## Route 4: Can a Mac Studio Cluster Do It? Apple Silicon is the obvious candidate for a memory-hungry model, because a 512GB M3 Ultra holds far more usable memory than any single GPU. One is not enough. To reach the roughly 2TB of aggregate memory K3 needs, you are looking at four or five 512GB Mac Studios, at roughly $9,500 each, wired together. That is where the plan meets physics, and the physics is latency, not weight bandwidth. Cluster tooling like EXO or mlx.distributed shards the model layer-wise, so each machine holds its own layers and their experts resident; what crosses the link is a small activation vector at each pipeline boundary, on every token. Thunderbolt 5 delivers roughly 6-8 GB/s in real transfers against the M3 Ultra's 819 GB/s of local memory bandwidth, about a hundred times slower, and the round trips serialize across four or five hops. The cost is the stalls between machines, not moving the weights themselves. Can it load? Probably, above about 2TB of usable aggregate memory. Is it a practical coding assistant? Nobody has published a K3 run on a Mac cluster yet. The closest datapoint is a prior-generation two-Mac rig with 1TB combined that reported about 15 tokens per second on other large models, plus an owner's own projection of a heavily-quantized K3 at a similar rate with, in their words, "abysmal" prefill. Treat 15 tokens per second as an analogy, not a measurement. For most people, a Mac cluster is a fascinating experiment, not a workstation. ## Route 5: Just Use a Provider For the overwhelming majority of use cases, the way to run Kimi K3 "locally" is to point your local tools at a hosted endpoint and keep working. K3 had been on [OpenRouter](https://openrouter.ai/moonshotai/kimi-k3) through Moonshot's own API since mid-July; within days of the weights going public, independent providers like Together AI and Modal were serving it there too. You add one API key to Cursor, VS Code, or your own scripts, and K3 answers from someone else's cluster. Speed varies a lot by host: as of August 2026 the range across the listed providers runs roughly 37 tokens per second at the slow end to about 148 at the fast end on Artificial Analysis' median, and Moonshot's own endpoint (about 39 tokens per second) now sits in the lower-middle of the pack rather than dead last. It is worth being precise about what open weights change here, because the popular version overstates it. The argument that held up best in launch-week discussion is that the payoff is not laptop inference, which almost nobody can do, but the option of many independent hosts. That has driven prices down for smaller open models. On K3 that pressure is only starting to show: as our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) breakdown documents, most third-party hosts still match Moonshot at $3 and $15, but as of August 2026 a few (Morph in fp4, DigitalOcean, DeepInfra) have begun undercutting to about $2.80 in and $14 out. The pool of operators who can host a 2.8-trillion-parameter model at all is small, so expect downward pressure to arrive slower than it did for 70B-class weights. What you get today is durability, since a version cannot be silently retired, and the option of data residency for teams that need it. Moonshot's API runs $3.00 per million input tokens, $0.30 on a cache hit, and $15.00 per million output. ## Is Self-Hosting Kimi K3 Worth It? For most people, no. Two conditions make it worth the trouble, and you need at least one: data you are legally required to keep in-house, or an existing GPU cluster that is already paid for. Absent both, the economics turn against you fast. The clearest way to see it is per-token cost. Practitioners running the numbers land between roughly $80 and $200 per million tokens on electricity alone, before hardware: a heavy-offload rig producing tokens slowly, on the order of 1 token per second, while drawing 2-5kW at typical US rates near $0.10-$0.15 per kWh. A faster decode rate lowers that, but you are still paying for hardware the API does not charge you for. Moonshot's API and its third-party hosts sit well below that. The exact figure swings with your workload, decode-bound jobs land cheaper and prefill-bound jobs land at the high end, but running K3 at home to save money is, for most workloads, more expensive than the API once you count the hardware you also had to buy. So the two legitimate self-host cases are narrow. If you have real GPU infrastructure and a data-custody reason, run K3 non-interactively, as an overnight batch job where the prefill wall does not matter. And split the decision by scale before you spec anything: a single-user batch setup is a $5,000 to $65,000 problem (current DDR prices push realistic 1TB+ builds toward the upper half), while serving a team with full context and real concurrency needs around 3TB of memory across the GPUs, for weights plus KV cache, and runs $300,000 to several million. There is also a harder question than "how" that a fixed budget forces: at a given amount of memory, do you run K3 at a punishing 2-bit quant, or a smaller frontier model at a comfortable 4- or 8-bit on the same box? The community has not settled it. One camp calls a 2-bit K3 a lobotomy; another argues its 104B active parameters should hold up better than most models at Q2; nobody had published memory, speed, and quality numbers side by side at the time of writing. Treat it as an open question, and if the answer is not obviously K3, the smaller frontier models in our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking will give you far more usable capability per dollar of hardware. The patient option the community keeps raising is to wait. Memory prices are elevated enough that builders call it "RAMageddon," with enterprise acquisition costs up an estimated 50-80% since late 2025 and some deferring rigs to 2028 or later, and a distilled 100B to 600B version of K3, runnable on far less hardware, is the outcome many expect. ## Frequently Asked Questions ### Can I run Kimi K3 on a Mac? Not on a single Mac. Even a 512GB M3 Ultra falls short of the 1.56TB the weights need. A cluster of four or five 512GB Mac Studios can hold the model, but the Thunderbolt link between them is about a hundred times slower than each machine's internal memory, so the interconnect sets the speed. Nobody has published a K3 run on a Mac cluster; the closest analogy, a prior-generation rig on other large models, sat around 15 tokens per second. It loads. Whether it is usable is another question. ### Do Ollama and llama.cpp support Kimi K3 yet? Not in a released build, as of August 2026. Hugging Face now lists roughly 29 community quantizations, but llama.cpp cannot read K3's architecture in any released build; support is still an open pull request (#26185), and the GGUF repos require compiling from that branch. Ollama's cloud tag routes to Ollama's servers rather than running locally. The working local serving paths today are vLLM and SGLang, both of which shipped day-0 support (SGLang on July 27, 2026), plus TokenSpeed. ### How big is the Kimi K3 download? The official repository is 1.56TB across 96 weight shards, in native MXFP4 4-bit precision. The widely repeated 594GB figure does not reconcile with 2.8 trillion parameters at 4-bit. Because K3 was quantization-aware-trained at 4-bit, there is no free halving from the usual 16-bit-to-4-bit step. Sub-4-bit community quants are smaller (a 2-bit build still runs north of 800GB), but they trade quality for the reduction rather than compressing a bloated original. ### Can I use Kimi K3 commercially if I self-host it? Yes, with conditions. K3 ships under the custom Kimi K3 License, which permits commercial use but is not OSI open source. If your model-as-a-service revenue passes $20 million over any 12 months you must negotiate a separate agreement with Moonshot, and a product above 100 million monthly users or $20 million in monthly revenue must display "Kimi K3" in its interface. Internal use carries no such obligation. Check the model card's license before you build. ### What is the cheapest way to run Kimi K3? A hosted provider, by a wide margin. Moonshot's API starts at $3 per million input tokens, and OpenRouter lists competing providers. Self-hosting on your own electricity can run $80 to $200 per million tokens on a slow rig once you count power, before hardware. Cheap local inference of a 2.8T model is not currently a thing. ## What K3's Launch Says About AI Search The hard question is the one we opened with. Any assistant answering from training data alone will keep getting K3 wrong for months, because a model cannot know about a release that happened after its cutoff, and retrieval only saves it if the engine can reach the page. That is not a trivia problem. It is the same mechanism that decides whether AI search engines describe your own product, pricing, and latest launch correctly, or repeat something stale. The self-inflicted half of that is worth ruling out first: whether AI crawlers can reach your content at all. Our [AI readiness checker](https://geotoolbox.ai/tools/ai-readiness) scores that in one request, so a fixable reachability gap is not what is keeping you out of the answer. ## Sources - Kimi K3 model card - Moonshot AI / Hugging Face - `huggingface.co/moonshotai/Kimi-K3` - Kimi K3 day-0 support - vLLM Blog, July 2026 - `vllm.ai/blog/2026-07-27-k3` - Kimi K3: Open Frontier Intelligence (technical report) - arXiv - `arxiv.org/abs/2607.24653` - Kimi K3 provider listing - OpenRouter - `openrouter.ai/moonshotai/kimi-k3` - Kimi K3 benchmarks, pricing, hardware requirements, and self-hosting - Northflank - `northflank.com/blog/what-is-kimi-k3-self-hosting` - Kimi K3 provider output-speed benchmarks - Artificial Analysis, July 2026 - `artificialanalysis.ai/models/kimi-k3` - deltafin (expert-streaming Kimi K3 on Apple Silicon) - GitHub - `github.com/gavamedia/deltafin` - Practitioner reports - r/LocalLLaMA, r/LocalLLM, r/kimi launch-week threads, July 2026 --- ## Kimi K3 vs Claude: Coding, Cost, and Reasoning Compared > Kimi K3 vs Claude, compared honestly: which wins on frontend coding, reasoning, speed, and real cost per task across Fable 5, Opus 5, and GPT-5.6 Sol. - Canonical: https://geotoolbox.ai/blog/kimi-k3-vs-claude - Published: 2026-07-26 · Updated: 2026-08-18 Ask Claude Fable 5 whether Kimi K3 is better than it, with web search switched off, and it tells you Kimi K3 does not exist. Ask Google's Gemini 3.1 Pro the same question and it goes further: it calls "Claude Fable 5" a hallucinated model name and insists neither model is real. Both are real, and both are current models as of July 2026. The assistants most people rely on cannot see either one yet. So here is the comparison those models cannot give you, current as of July 2026: Kimi K3 vs Claude, head to head on coding, reasoning, speed, and what each actually costs. There is no single winner; the honest answer depends on the job. We flag which numbers are independently measured and which come from the labs themselves, because on models this new that distinction is most of the story. ## Kimi K3 vs Claude: The Short Answer **Neither wins outright. Kimi K3 leads on frontend and visual coding, on raw cost per task, and on being open. Claude leads on general reasoning, speed, long-horizon work, and the reliability that matters when a wrong answer is expensive.** The teams getting the most out of both are not switching from one to the other. They route by task and keep both. That framing matters because most "Kimi K3 vs Claude" takes are really "should I cancel my Claude subscription," which is the wrong question. K3 is a coding and agent specialist that happens to be cheap per token. Claude, in its [Opus 5](https://geotoolbox.ai/blog/claude-opus-5) and Fable 5 configurations, is a broader reasoner that happens to be fast. You choose per workload, not per loyalty.
DimensionWinnerWhy
Frontend / visual codingKimi K3First open model to top the Frontend Code Arena, ahead of Fable 5
General intelligenceClaudeOpus 5 leads the Artificial Analysis Intelligence Index; K3 trails all three
Hard reasoningClaudeFable 5 and Opus 5 clear K3 by roughly ten points on Humanity's Last Exam
Cost per taskKimi K3Around $0.95 per Intelligence Index task versus $2.03 for Opus 5
SpeedClaudeFable 5 outputs roughly twice as fast; Opus 5 starts a task in about a third of the time
Open weightsKimi K3Open-weight model (weights public since July 27, custom license); Claude is closed and API-only
Refusal-sensitive workKimi K3 (reported)Developers report fewer blanket refusals on security- and medical-adjacent work
Each row below has evidence behind it. The one almost every comparison skips comes first: which Claude are you even talking about? ## Which Claude Are You Comparing Against? "Claude" is not one model, and this is where most K3 comparisons go wrong. A benchmark that says "Claude scored 53" is useless unless you know whether it means Fable 5, Opus 5, or Sonnet 5, because those three sit at different points on both the capability and the price curve. K3 beats one of them on a given task and loses to another on the same task. Three configurations matter for this comparison. **Fable 5** is Anthropic's top-tier model for long-running autonomous agents, and its most expensive. **Opus 5**, released July 24, 2026, tops the independent intelligence leaderboard and is the sensible high end for reasoning. **Sonnet 5** is the balanced production default, and the one whose sticker price lands closest to K3. There is also a legacy Opus 4.8 you will still see in older benchmark tables.
Claude modelBuilt forAPI price (input / output per 1M)Note
Fable 5Long-running autonomous agents$10 / $50Top tier and priciest; overkill for jobs Opus can do
Opus 5High-end reasoning and coding$5 / $25Current independent intelligence leader
Sonnet 5Balanced production default$2 / $10 (permanent rate)Closest Claude to K3 on price
Opus 4.8 (legacy)Prior-gen reasoning$5 / $25Superseded by Opus 5; common in old tables
When a headline says K3 "beats Claude," it almost always means it beat Opus 4.8 or Sonnet on a coding leaderboard. It rarely means it beat Fable 5 or Opus 5 on reasoning. Keeping the specific model attached to every number, which the tables below do, is the difference between a real comparison and a marketing screenshot. Our [Claude pricing guide](https://geotoolbox.ai/blog/claude-pricing) breaks down the full lineup and the subscription tiers that bundle each model. ## The Benchmarks: Who Wins What On a two-week-old model, sort every number by who measured it before you trust it. Independent evaluators run the same harness across models; the labs run their own harness, at maximum effort, on the tasks that flatter them. Both appear below, labeled, because the split is the story. The independent picture is consistent. On the [Artificial Analysis](https://artificialanalysis.ai/models/kimi-k3) Intelligence Index, K3 scores 57 and sits seventh of 190 models, behind Opus 5 at 61 and Fable 5 at 60. But on the [Frontend Code Arena](https://arena.ai/leaderboard), a human-preference vote on generated interfaces that leans toward frontend and visual tasks, K3 ranks first, the first open-weight model to lead it, and it beats Fable 5 head to head. Read those two results together: K3 is a coding and interface specialist at the frontier of its lane, and only competitive outside it.
![Who-wins-what scorecard comparing Kimi K3, Claude, and GPT-5.6 Sol across six dimensions: Kimi K3 leads frontend coding (Frontend Code Arena Elo 1,682), cost per task ($0.95), and open weights, while Claude leads general intelligence (Opus 5 at 61), hard reasoning (53 on Humanity's Last Exam), and output speed (Fable 5 at 71 tokens per second).](/blog/kimi-k3-vs-claude/kimi-k3-vs-claude-benchmarks.png)
Independently measured results tell a different story from the vendor tables. K3 owns frontend and cost; Claude owns reasoning and speed.
BenchmarkKimi K3ClaudeGPT-5.6 SolMeasured by
Frontend Code Arena (Elo)1,682 (#1)1,630 (Fable 5)1,625Independent (Arena)
AA Intelligence Index57 (#7 of 190)61 (Opus 5), 60 (Fable 5)59Independent (Artificial Analysis)
Humanity's Last Exam (reasoning, no tools)43.553 (Fable 5 and Opus 5)n/aMixed (K3 vendor; Claude independent)
Long-horizon work (GDPval Elo)1,6861,861 (Opus 5)1,735Independent (Artificial Analysis)
Terminal-Bench 2.188.384.6 (Fable 5); 89 (Opus 5)88.8Mixed (Opus 5 independent; rest vendor)
GPQA Diamond93.592.6 (Fable 5)94.1Vendor-reported (Moonshot table)
Accuracy (AA-Omniscience)46%61% (Fable 5)59%Independent (Artificial Analysis)
One nuance sits behind that accuracy row. Fable 5 is the most accurate model in the set on Artificial Analysis's Omniscience test, answering 61% of hard questions correctly to K3's 46%. But accuracy and honesty are not the same thing: K3 lifted both its accuracy and its fabrication rate over the older K2.6, so it now gets more right and invents more of what it gets wrong. Verify anything that matters, whichever model produced it. For the full K3 architecture and its launch-day benchmark caveats, see our [Kimi K3 explainer](https://geotoolbox.ai/blog/what-is-kimi-k3). ## Price: Why the Sticker Is Misleading On paper this is the easiest win in the comparison. Kimi K3 charges $3 per million input tokens and $15 per million output, with cache hits dropping input to $0.30. That output rate runs 50% above Sonnet 5's permanent $10 rate, undercuts Opus 5, and is a third of Fable 5. And on Artificial Analysis's controlled cost-per-task measure, K3 is the cheapest model in the set at about $0.95, under Opus 5 at $2.03 and Fable 5 at $2.75. If you stop reading here, K3 looks like a straight price cut.
ModelInput / 1MOutput / 1MCached inputCost per task (AA, max)
Kimi K3$3.00$15.00$0.30$0.95
Claude Sonnet 5$2.00 (permanent rate)$10.00~$0.20$1.53
Claude Opus 5$5.00$25.00$0.50$2.03
Claude Fable 5$10.00$50.00~$1.00$2.75
Two things complicate the clean number. The first is verbosity. K3 reasons on every request, and with its reasoning effort locked to the single "max" setting on Moonshot's API and OpenRouter alike, it meters far more output than a comparable Claude call: Artificial Analysis clocked K3 at 130M output tokens to run its index, against a 63M median. On an agent job with retries and a growing history, that extra output lands on the $15 line, which is why first-week users on Kimi's coding plans repeatedly report a single task eating a startling share of their quota. The second is a comparison error worth naming, because so many of the "K3 is 10x cheaper" posts are built on it. **People compare K3's per-token API price to their flat monthly Claude subscription and conclude K3 is far cheaper. Those are different products.** A $20 Claude Pro plan is not billed per token; a fair comparison is API against API (where K3 undercuts Opus and Fable but sits near Sonnet), or subscription against subscription (where Kimi's coding plans have their own tiers and their own quota complaints). Our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) guide covers the token rates and tiers, and the [Claude pricing guide](https://geotoolbox.ai/blog/claude-pricing) does the same on the Anthropic side. Compare like with like and the gap narrows from "10x" to "real but workload-dependent." ## Kimi Code vs Claude Code: The Agentic Head-to-Head For most developers the real comparison is not the raw model but the coding agent wrapped around it. Here the two are closer than the leaderboards suggest, and the deciding factors are speed, quota clarity, and whether you can get access at all. You do not have to choose the tool to test the model. K3 is available through OpenRouter and through Moonshot's own API, so you can point [Claude Code](https://geotoolbox.ai/blog/what-is-claude-code) at K3 and run it inside the harness you already know, and many people do exactly that to A/B the two on the same task. Moonshot also ships its own agent, Kimi Code, with its own subscription tiers. Either way, the first thing you notice is speed. K3 outputs around 32 tokens per second, roughly half Fable 5's 71, and it is slow to start: close to three minutes to the first token where Opus 5 takes about one. On an interactive coding loop, that wait is felt on every turn, and it partly cancels the price advantage. Context is a wash: K3 carries a one-million-token window, and Claude Opus 5 matches it, so neither wins on raw context. Three practical traps come up repeatedly in first-week reports, and none are about capability: 1. **Capacity, not quality, is the current blocker.** Demand outran Moonshot's serving capacity at launch, so through late July 2026 Kimi's coding plans sold out and new sign-ups hit a waitlist. The best model in the world is no use if you cannot buy access this week. 2. **Silent misconfiguration ruins the test.** Pointing a tool at the wrong model ID, or a tier that caps context below the advertised window, quietly hobbles K3 before the comparison even starts. Confirm the exact model string and your context ceiling first. 3. **The quotas are opaque.** Kimi's coding plans report usage as a percentage rather than tokens, so you cannot easily see what a task cost, which is why some users route their subscription through a gateway just to get visibility. Claude Code and its context handling are more legible, and our notes on the [Claude Code context window](https://geotoolbox.ai/blog/claude-code-context-window) and [cutting Claude Code token costs](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) apply directly if bill control is the goal. Across serious testers, the pattern is consistent: they keep Claude Code as the daily driver and reach for K3 on the specific jobs, dense frontends and long-context passes, where it measurably pulls ahead. That is routing in practice. ## Open Weights vs Closed: What You Actually Get This is the one axis where the comparison is not close, and also the one most people misread. Kimi K3 is an [open-weight](https://geotoolbox.ai/glossary/open-weights) model: Moonshot published the weights on July 27, 2026 to its Hugging Face repo, under a custom Kimi K3 License that permits commercial use with conditions at very large scale. Claude is closed. You reach it only through Anthropic's API or apps, and you never hold the model. The trap is assuming "open weights" means "I can run this myself." For almost everyone, you cannot. At 2.8 trillion parameters, the repository is about 1.5TB in K3's native 4-bit MXFP4 format, and you need roughly 1.4TB of GPU memory just to load it, so Moonshot recommends a supernode of 64 or more accelerators. That is data-center hardware, not a workstation with a good graphics card. Community quantizations for llama.cpp and Ollama did land quickly, but they shrink the hardware bill without erasing it. Our [run an LLM locally](https://geotoolbox.ai/blog/run-llm-locally) guide and our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking both include a self-host reality check for exactly this class of model. So what does open buy you against Claude, if not laptop inference? The sharpest version practitioners land on is that a model this large is legally open but operationally closed to anyone without a cluster, and the payoff is not that you personally run it. It is that anyone can, which creates a market: because the weights are public, competing providers host K3 and drive the API price down, the way many independent hosts already do for models like GLM. On top of that, you can keep data in your own jurisdiction if you self-host, fine-tune on your own domain, and stay insulated from a provider changing terms, prices, or pulling a model. That residency point comes with a catch worth stating plainly: the hosted K3 API sends your prompts to Moonshot's servers in China, so the data benefit only holds if you run the weights yourself. If none of these apply, "open" is a transparency benefit rather than a practical one, and Claude's closed, hosted convenience is the better trade. The distinction, and why it is not the same as "open source," is covered in our [open weights vs open source](https://geotoolbox.ai/blog/open-weights-vs-open-source) explainer. ## Adding GPT-5.6 Sol: The Three-Way Most head-to-head videos are actually three-way: Kimi K3 versus Claude Fable 5 versus [GPT-5.6 Sol](https://geotoolbox.ai/blog/gpt-5-6), all building the same app on camera. It is worth knowing where OpenAI's model lands, because it changes the decision at the margins. Sol sits between the other two on most measures. It leads the set on GPQA Diamond at 94.1, edges ahead of K3 on the Intelligence Index at 59, and roughly matches K3 on Terminal-Bench, all while costing about $1.04 per task, close to K3 and under both Opus 5 and Fable 5. Where it clearly loses is the specific thing K3 was built for: on the Frontend Code Arena it trails both K3 and Fable 5 for building interfaces. The three-way verdict is cleaner than it sounds. For raw frontend and interface work, K3 is the pick. For careful reasoning and the hardest general problems, Claude Opus 5 or Fable 5. For a fast, cheap, broadly capable middle that rarely embarrasses itself, GPT-5.6 Sol is the safe default, and it is why many teams keep it as the everyday model and treat both K3 and Claude as specialists they call in. If you are weighing the wider field, our [ChatGPT alternatives](https://geotoolbox.ai/blog/chatgpt-alternatives) rundown puts all three in context. ## Which Should You Use? A Decision Framework Here is the direct answer the title asks for: **for building dense frontends, agent-heavy coding, and cost-sensitive high-volume work, Kimi K3 is worth adopting. For careful reasoning, unfamiliar-repo surgery, latency-sensitive interactive work, and anything where a confident wrong answer is expensive, stay on Claude.** Neither replaces the other; match the model to the job.
If your job is...Reach forBecause
Building or restyling a frontend / UIKimi K3Ranks first for interface generation, beating Fable 5
High-volume coding on a tight budgetKimi K3Lowest cost per task, if you can tolerate the latency
Data must stay in your jurisdictionKimi K3, self-hostedOnly if you run the weights yourself; the hosted API sends data to China
Hard reasoning or a wrong answer is costlyClaude Opus 5 / Fable 5Leads reasoning benchmarks and is more accurate
Interactive, latency-sensitive workClaudeRoughly twice the output speed and a much shorter wait to first token
Refusals are blocking legitimate workKimi K3 (reported)Developers report fewer blanket refusals; not independently benchmarked
You want one dependable everyday modelGPT-5.6 Sol or Claude Sonnet 5Broadly capable, cheap, rarely a surprise
Do not switch on launch-week benchmarks. Pilot instead: run K3 and your current Claude model on one real, measurable task, look at accepted output and how much supervision each needed, and keep whichever leaves less total friction. For where K3 fits among the other open Chinese models, our [Chinese AI models comparison](https://geotoolbox.ai/blog/chinese-ai-models-compared) is the wider map. ## What This Means for Your AI Visibility Return to where we started. With web search off, both Claude Fable 5 and Gemini denied that Kimi K3 exists, and one insisted a real Claude model was fictional. Two of the most capable systems on earth, blind to a launch that led the AI news cycle for a week, because their training predates it. That lag is not a quirk of new model launches. It is how every assistant treats anything newer than its training, including your business. A product you shipped last month, a rebrand, a corrected fact about what you do, all of it is invisible to deployed assistants until training catches up or they read it live on the web. And now that K3's weights are public, the model gets fine-tuned into a long tail of downstream tools you will never see, each answering questions about your market from whatever it can find about you. That makes reachability the lever, not the model. The systems that update faster than training do so by fetching live pages, so if AI crawlers cannot reach and parse your site, you are absent from the layer that stays current, and correct answers about you never get a chance to form. This is the core of [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) and of [getting cited by AI](https://geotoolbox.ai/blog/what-is-geo): the businesses that surface well are the ones a model can find, read, and trust without tripping over contradictions. You cannot control what Kimi K3 or the next model learns about you. You can control whether it can reach you at all. Run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether AI crawlers can actually fetch and parse your site, and close the gaps before the next launch makes the question urgent. ## Frequently Asked Questions ### Is Kimi K3 better than Claude? For frontend and interface coding, yes: K3 ranks first on the human-preference Frontend Code Arena, ahead of Claude Fable 5. For general reasoning and the hardest problems, no: Claude Opus 5 and Fable 5 lead the independent Intelligence Index and score about ten points higher on Humanity's Last Exam. K3 is a specialist that wins its lane, not an across-the-board replacement. ### Is Kimi K3 cheaper than Claude? Against Opus 5 and Fable 5, yes: K3's per-token rates undercut both, and on Artificial Analysis's cost-per-task measure K3 is the cheapest model in the set at about $0.95, below Opus 5's $2.03. Sonnet 5 is the exception, cheaper per token today, though K3 is still cheaper per task. And K3's verbose reasoning meters more output on real jobs, so comparing its API price to a flat Claude subscription is not like-for-like. ### Can I use Kimi K3 in Claude Code? Yes. K3 is served through OpenRouter and Moonshot's own API, so you can point Claude Code at it and run it in the harness you already use. Many developers do this to test K3 against their usual Claude model on the same task before deciding anything. ### Is Kimi K3 faster than Claude? No, it is noticeably slower. K3 outputs around 32 tokens per second, about half Fable 5's 71, and it is slow to start, taking close to three minutes to its first token where Opus 5 takes about one. On interactive coding, that lag shows on every turn. ### Are Kimi K3's weights available yet? Yes. Moonshot published the open weights on July 27, 2026 to its Hugging Face repo, under a custom Kimi K3 License. But the model is far too large to run on consumer hardware, about 1.5TB to download and needing a large multi-node GPU cluster to serve (Moonshot recommends 64-plus accelerators), so hosting through an API remains the practical route for most. ### Does Kimi K3 hallucinate more than Claude? It cuts both ways. On Artificial Analysis's Omniscience test, Fable 5 is the most accurate model, answering 61% of hard questions correctly to K3's 46%. But K3's fabrication rate has risen as its accuracy climbed over the older K2.6, so it gets more right and makes up more of what it gets wrong. Both invent enough that anything important should be verified. ## Sources - Kimi K3 vs Claude Opus 4.8 - Intelligence, cost, and speed comparison - Artificial Analysis, July 2026 - `artificialanalysis.ai/models/comparisons/kimi-k3-vs-claude-opus-4-8` - Kimi K3 - Intelligence, Performance & Price Analysis - Artificial Analysis, July 2026 - `artificialanalysis.ai/models/kimi-k3` - Claude Opus 5 - Intelligence Index and cost per task - Artificial Analysis, July 2026 - `artificialanalysis.ai/articles/opus-5` - Frontend Code Arena leaderboard - Arena, July 2026 - `arena.ai/leaderboard` - Claude API pricing and model rate card - Anthropic, July 2026 - `platform.claude.com/docs/en/about-claude/pricing` - Kimi K3 - API pricing and model card - OpenRouter, July 2026 - `openrouter.ai/moonshotai/kimi-k3` - Kimi-K3 model repository (open-weights release) - Hugging Face / Moonshot AI, July 2026 - `huggingface.co/moonshotai/Kimi-K3` - Kimi K3, and what we can still learn from the pelican benchmark - Simon Willison, July 2026 - `simonwillison.net/2026/Jul/16/kimi-k3` --- ## OKF vs RAG: Does Google's Open Knowledge Format Replace RAG? > Google's Open Knowledge Format (OKF) is being called a RAG killer. It isn't. What OKF v0.2 actually does, where it beats RAG, and how the two work together. - Canonical: https://geotoolbox.ai/blog/okf-vs-rag - Published: 2026-07-26 · Updated: 2026-07-26 Search for "OKF vs RAG" and you get two stories. One says Google just killed RAG. The other says the two have nothing to do with each other. Both are wrong. The Open Knowledge Format (OKF) and retrieval-augmented generation (RAG) solve different problems at different layers. The useful question is not which one wins but where each belongs, including the part almost every explainer still skips: what changed in OKF v0.2. There is a fitting demonstration of why curated knowledge matters. Ask an AI model with no web access what OKF is: in a test we ran, Gemini invented a wrong definition while Claude said it did not know. OKF launched in mid-2026, after both models' training cutoffs, so only a model wired to live retrieval gets it right. Which is, more or less, the whole argument for writing your knowledge down where machines can read it.
![OKF plus RAG hybrid architecture: a router sends curated known-knowns to OKF and the unstructured long tail to RAG, feeding one grounded answer.](/blog/okf-vs-rag/okf-rag-hybrid-architecture.png)
One agent, two knowledge layers: OKF owns the curated core, RAG keeps the long tail.
## What Is the Open Knowledge Format (OKF)? **The Open Knowledge Format (OKF) is an open, vendor-neutral specification from Google Cloud for handing curated knowledge to AI agents as plain Markdown files.** [Google Cloud introduced it](https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing) on June 12, 2026 as a way to package the facts an organization already knows and trusts, so an agent can read them directly instead of re-deriving them from raw documents every time. The format itself is deliberately boring, which is the point. An OKF bundle is a directory tree of Markdown files, one concept per file, each carrying a small block of YAML frontmatter. A concept can be a metric definition, a database schema, a business rule, an API, or a policy. Files link to each other with ordinary Markdown links, so a cross-linked bundle reads as a lightweight [knowledge graph](https://geotoolbox.ai/glossary/knowledge-graph) an agent can walk, not a pile of disconnected pages. Per the [OKF specification](https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md), the only always-required field is `type`. A concept carrying nothing but `type` is already valid. Recommended fields (`title`, `description`, `resource`, `tags`) add context, and two reserved filenames do the structural work: `index.md` lists a directory for progressive disclosure, and `log.md` records history. The idea is not new. It formalizes the "LLM wiki" pattern [Andrej Karpathy floated in April 2026](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f): a persistent set of Markdown files an agent maintains, compiled once and kept current rather than reconstructed on every query. OKF gives that pattern a shared contract so any agent or vendor can read the same bundle. Two caveats. OKF ships inside Google Cloud's Knowledge Catalog under the Apache 2.0 license, but the repository carries a plain "not an official Google product" disclaimer. And it is early: Google published three sample bundles (GA4 e-commerce, Stack Overflow, and Bitcoin datasets) to seed adoption, but the ecosystem around it is weeks old. ## What Is RAG, and Where Does It Break Down? [Retrieval-augmented generation (RAG)](https://geotoolbox.ai/blog/what-is-rag) is the standard way to connect a language model to knowledge it was not trained on. You ingest a large corpus, split each document into chunks, embed those chunks as vectors, and at query time retrieve the ones that sit closest to the question before the model writes its answer. For a big, messy, fast-changing corpus, this works well and it is not going anywhere. The trouble starts when RAG is asked to serve knowledge that is small, stable, and used constantly, like how your company defines an active user. The bare RAG pattern does not store that definition as a canonical fact. It re-derives it from chunks at query time, and that re-derivation has four recurring failure modes. The first is [chunking](https://geotoolbox.ai/blog/content-chunking). Splitting a document into fixed windows cuts across logical boundaries, so a chunk read in isolation can look irrelevant even when it holds the answer. The second is that [vector similarity](https://geotoolbox.ai/blog/vector-embeddings) rewards text that looks relevant, not text that is correct. A passage sitting near your query in embedding space can still be the wrong one, and on a precise technical fact "near enough" is wrong. The third is that these failures are silent. RAG rarely errors out. It returns a fluent, plausible, subtly wrong answer, which is the most expensive kind on a high-stakes fact. The fourth is structure: in a plain vector setup, the model receives a bag of chunks with no stated relationships between them, so any connection between two facts has to be inferred rather than read. (Graph and hierarchical RAG variants add some of this back, at the cost of more machinery.) None of this makes RAG bad. It makes RAG the wrong tool for the slice of knowledge that is already curated and needs to be exactly right every time. That slice is the gap OKF is built for. ## OKF vs RAG: The Core Difference The cleanest way to hold the two in your head: RAG is a retrieval pattern, and OKF is a knowledge format. They operate at different layers, which is exactly why "OKF vs RAG" is the wrong framing for most decisions. One decides how an agent finds information; the other decides how that information is written down in the first place. RAG treats knowledge as a large body of text to be searched at query time. OKF treats a defined set of concepts as durable facts to be read on demand. Where RAG reconstructs meaning from whichever chunks rank highest, an OKF bundle hands the agent a concept someone already got right, with its relationships to other concepts stated rather than inferred.
 OKF (Open Knowledge Format)RAG (Retrieval-Augmented Generation)
Knowledge typeCurated known-knowns you can nameLarge, unstructured, "it's in here somewhere"
How the agent gets itReads a concept file directlyRetrieves and reranks chunks at query time
StructureExplicit graph via Markdown linksA bag of chunks, relationships inferred
Provenance & trustStandardized fields (sources, verified)Per-implementation metadata, not standardized
FreshnessDeclared with stale_afterDepends on re-indexing the corpus
Best forStable, high-stakes factsLong-tail, exploratory search
Read the provenance and freshness rows again. That is where OKF v0.2 standardizes something [retrieval](https://geotoolbox.ai/glossary/retrieval-augmented-generation) leaves to each implementation. ## The v0.2 Trust Layer: What OKF Standardizes That RAG Doesn't The version that matters is 0.2. It adds what amounts to a trust layer to the spec, and it is the reason the versus framing keeps missing the point. Almost every OKF-vs-RAG explainer online still describes v0.1 and skips this entirely. RAG can tell you what a passage says. The bare pattern does not standardize who wrote it, whether anyone checked it, or whether it is still true. Individual RAG systems can bolt those on as chunk metadata, but every team invents its own schema. OKF v0.2 writes them into the frontmatter as a shared contract, through the optional field families below.
FamilyKey fieldsWhat it answers
Provenancesources, usage_countWhere did this come from, and how much is it relied on?
Trustgenerated, verifiedWho produced it, and who confirmed it?
Lifecyclestatus, stale_afterIs it current, and when should it expire?
Attestationruntime, computation, attesterWas this value actually computed, and by what?
The `verified` field is the centerpiece. It turns a flat file into a claim with a trust tier: unverified, machine-confirmed, or human-reviewed (marked with a `human:` actor). An agent can then treat a human-reviewed metric definition differently from an unverified note. A vector store can carry that in metadata too, but OKF makes it a defined field every reader interprets the same way, instead of a per-pipeline convention. This is also why v0.2 restructured v0.1's flat `timestamp` into `generated: { by, at }`: a fact now records who produced it and when, not just a bare date. A concept with the v0.2 fields looks like this: ```yaml --- type: metric title: Weekly Active Users verified: - { by: "human:data-team", at: 2026-07-20 } stale_after: 2026-10-20 sources: - id: wau-sql resource: /queries/weekly-active-users.sql title: WAU definition query --- ``` `stale_after` is the quiet one. The bare RAG pattern has no built-in notion of a fact expiring; unless you add time-to-live logic yourself, a stale chunk sits in the index looking exactly as retrievable as a fresh one. OKF lets the knowledge itself declare an expiry date, which pushes staleness from an invisible failure into a field you can query and act on. Combined with the `sources` block, this gives a bundle a kind of [grounding](https://geotoolbox.ai/glossary/grounding) and accountability that a retrieval index does not carry by default. ## Does OKF Replace RAG? The Hybrid Split No. The headlines calling OKF a RAG killer are selling a clean story that the people actually building with it do not repeat. The working consensus is a hybrid, and the division is lopsided. The small share of facts that are stable, critical, and reused constantly (metric definitions, business rules, schema, regulated facts) lives in OKF, where it is written once and read directly. The large majority, the long tail of unstructured PDFs, tickets, logs, and changing documents, stays in RAG, where retrieval is the only sane option. A router in front of the agent decides which path a query takes: a lightweight classifier, or the agent itself, sends known-knowns to the bundle and everything else to the vector store. The bundle reaches the model the way any file does, through a file-read tool or an MCP resource, and `index.md` lets the agent scan a directory before pulling the one concept it needs rather than loading the whole bundle into context. These two are not even mutually exclusive at the storage layer. An OKF bundle is clean, well-structured Markdown, which happens to be close to ideal input for a RAG chunking pipeline: point retrieval at the bundle and you get better chunks than raw source documents would give. OKF improves the knowledge; RAG still handles [how it gets found](https://geotoolbox.ai/blog/query-fan-out) at scale. There is a minority position worth stating fairly. Some writers, [Alphamatch](https://www.alphamatch.ai/blog/google-open-knowledge-format-okf-vs-rag-2026) among them, argue OKF removes the need for RAG in many enterprise agent scenarios, since re-retrieving stable facts from chunks is pure waste. They have a point about that slice. But "replaces RAG for facts you have already curated" is a far narrower claim than "replaces RAG," and it collapses the moment your agent needs anything outside the curated set. So the plain answer to "is RAG obsolete in 2026" is that OKF displaces vector retrieval for the curated facts you can address directly, and leaves the long tail exactly where it was. ## When to Use OKF, RAG, or Both Set the labels aside and ask four questions about the knowledge itself: How big is it? How often does it change? How badly does the agent need to be right? And how often is it reused? Those answers point at a tool.
Your situationUseWhy
Metric definitions, business rules, schemasOKFSmall, stable, reused constantly, must be exact
Millions of support tickets, PDFs, logsRAGToo large and unstructured to curate by hand
Regulated facts the agent must never get wrongOKFThe verified tier records a human-reviewed audit trail
Fast-changing, exploratory research corpusRAGCuration cannot keep pace with the churn
An enterprise agent that needs bothHybridRoute known-knowns to OKF, the long tail to RAG
The rule underneath the table is simple: when you can name the fact and getting it wrong is expensive, curate it in OKF; when the knowledge is a haystack you search occasionally, leave it in RAG. Most real systems land on the last row, because production agents need both a handful of things they are certain about and a large body of things they can look up. One decision worth making early: what belongs in the curated core. The temptation is to over-curate and turn OKF into a second unmaintained wiki. The test is reuse under pressure. If an agent reads a fact on nearly every task and a wrong answer causes real damage, it earns a concept file. Everything else stays in retrieval. Cost cuts the same way. OKF avoids embedding spend, a vector database, and retrieval latency for the facts it holds, but loading many concept files burns context tokens and will not scale to thousands of them. That points back at the same split: a small curated core in OKF, the bulk in RAG. ## OKF vs Knowledge Graphs, Wikis, and llms.txt The most common objection to OKF is that it is nothing new. It is worth taking seriously, because it is half right. A folder of Markdown files with metadata is not a novel invention. What is new is the shared contract. Against a formal [knowledge graph](https://geotoolbox.ai/glossary/knowledge-graph), OKF trades power for maintainability. A triple store with a query language gives you richer traversal, but it also demands specialized tooling and people who know it. OKF is human-readable files anyone on the team can edit in a text editor and review in a Git pull request. You get the graph shape, through cross-links, without the graph database. Against a wiki, Obsidian vault, or a `CLAUDE.md` file, the difference is standardization. Those all work, but each one is bespoke. OKF fixes the frontmatter fields and the required `type`, so any agent or vendor that speaks OKF can read your bundle without custom parsing. Portability is the entire pitch, and it is why the "just a folder" dismissal understates it. Against [llms.txt](https://geotoolbox.ai/blog/llms-txt), the two are complementary, not competing. An llms.txt file is a pointer, a curated index that tells an agent where to look on your site. OKF is the structured knowledge the agent reads once it gets there. You can publish both, and for AI discoverability they stack cleanly with [schema markup](https://geotoolbox.ai/blog/schema-markup-for-ai) and [entity data](https://geotoolbox.ai/blog/entity-seo) as layers of the same job: making your knowledge legible to machines. So OKF is not a breakthrough in data modeling. It is a standard wrapped around a pattern that already worked, which is a less exciting claim and a more useful one. ## How to Build and Host an OKF Bundle Building a bundle is closer to writing documentation than to standing up a database. The steps are short. Start with a directory of Markdown files, one concept per file. Give each file YAML frontmatter with at least a `type`, then add the recommended `title` and `description`, and layer in the v0.2 `verified` and `stale_after` fields wherever a fact is high-stakes enough to warrant them. Cross-link related concepts with plain Markdown links so the agent can walk between them. Add an `index.md` at each level so an agent can read a directory listing before opening every file. The result is a tree like this: ```text okf/ index.md metrics/ weekly-active-users.md revenue.md policies/ refund-policy.md ``` Then host it. A bundle is just files, so the options are ordinary: commit it to a Git repository, ship it as a tarball, or serve it from your site under a path like `/okf/`. Nothing about it requires Vertex AI, a vector database, or a Google account, which is what "vendor-neutral" means in practice, at least by design: no other major vendor has adopted OKF yet. One caution before you publish. A bundle of metric definitions, business rules, and schemas is internal knowledge. Serve publicly only what you would want public, and keep sensitive concepts behind auth or off the public tree entirely. The most useful real example comes from GEO practitioner [Suganthan Mohanadasan](https://suganthan.com/blog/open-knowledge-format/), who converted his published writing into a live bundle and built a free generator that turns a site or sitemap into OKF without code. His live bundle, at the time of writing, runs 61 concepts, of which only 18 carry a `verified` entry and eight are already past their `stale_after` date. That is the trust layer working as intended: the maintenance signal is visible in the file instead of hidden in a stale index. Google's three sample bundles are another starting point if you would rather adapt than build from scratch. The one caveat to set expectations: as of this writing, no major crawler ingests OKF bundles automatically. Publishing one is a bet on where [agent-ready sites](https://geotoolbox.ai/blog/agent-ready-website) are heading, not a switch that lights up traffic today. ## OKF and AI Visibility: Should SEOs Adopt It Yet? Here is the verdict, since the heading asked: for internal agent knowledge, adopt it now. For public AI visibility, treat it as a low-cost bet, not a ranking tactic. Publishing an OKF bundle is a cheap, low-risk way to hand curated facts to any agent that reads your site, and it stacks neatly with the other machine-readability layers. GEO practitioners are already experimenting with it on that basis, [Marie Haynes](https://www.mariehaynes.com/okf) among them. But there is no demonstrated citation or ranking lift from doing it, and the skeptics on r/TechSEO are right that OKF was built for organizational knowledge, not search. The sharper worry is adoption: OKF risks becoming another Google standard nobody uses, in the same lineage as protocols Google has shipped and quietly abandoned. The data world has a long graveyard of formats that went nowhere. Weigh both honestly and the move is small. Building a bundle for your own agents pays off now, because it fixes real RAG failure modes on facts you cannot afford to get wrong. Publishing it publicly is speculative, which is fine as long as you treat it that way and do not expect rankings to move. For anyone already asking whether AI systems can read and cite their site, OKF is that same question pointed inward at agents instead of outward at search engines. It is the discipline we apply at geotoolbox when we [check whether AI crawlers can reach and parse a site](https://geotoolbox.ai/blog/how-to-get-cited-by-ai): the format is new, but the job, making your knowledge legible to machines, is not. ## The Bottom Line OKF is not the end of RAG. It is the curated-facts layer RAG never had, and the practical architecture uses both: OKF for the small set of things your agent must be certain about, RAG for the large set it looks up. The v0.2 trust layer, with provenance, verification tiers, and expiry, is the real advance and the part worth building around, not the "RAG is dead" headline. If you are weighing OKF because you want AI systems to read your knowledge correctly, that instinct is the same one that decides whether AI search cites you at all. You can check where your site stands with our free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness), which tells you whether agents and AI crawlers can actually reach and parse what you have published, before you decide what to structure next. ## Frequently Asked Questions ### Is RAG obsolete in 2026? No. OKF replaces RAG only for the slice of knowledge that is small, stable, and already curated. For the long tail of large, unstructured, changing content, retrieval is still the right tool, and most production stacks run both. ### Does OKF replace the vector database? Not in general. It removes the need to re-retrieve stable facts from a vector store, but anything outside your curated concept set still needs retrieval. In hybrid setups an OKF bundle can even feed the RAG pipeline as clean input rather than replacing it. ### Do I need Vertex AI or Google Cloud to use OKF? No. An OKF bundle is plain Markdown files with YAML frontmatter. It requires no SDK, no vector database, and no Google account. You can store it in any Git repository or serve it from your own site. ### Is OKF the same as llms.txt? No, and they work together. An llms.txt file is an index that points agents to content. OKF is the structured knowledge itself. One tells an agent where to look; the other is what it reads when it gets there. ### What's the difference between OKF and a knowledge graph? A formal knowledge graph stores facts as queryable triples and needs specialized tooling. OKF gets a graph shape from Markdown cross-links while staying human-readable and Git-friendly. You trade some query power for far lower maintenance cost. ### Is OKF an SEO ranking factor? There is no evidence it affects rankings. OKF was designed for internal agent knowledge, not search, and no major crawler ingests bundles automatically yet. Publishing one is a bet on agent readability, with no ranking payoff to expect. ## Sources - How the Open Knowledge Format can improve data sharing - Google Cloud (official OKF announcement, June 2026, and the three sample bundles) - `cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing` - okf/SPEC.md - GoogleCloudPlatform/knowledge-catalog, GitHub (the v0.2 specification: required fields, trust layer, Apache 2.0 license) - `github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/okf/SPEC.md` - llm-wiki - Andrej Karpathy, GitHub Gist (the LLM-wiki pattern OKF standardizes, April 2026) - `gist.github.com/karpathy/442a6bf555914893e9891c11519de94f` - The OKF from Google is a new layer for agents - Marie Haynes (GEO practitioner treating OKF as an agent-knowledge layer; early hands-on experiments) - `mariehaynes.com/okf` - Open Knowledge Format (OKF): Google's New Markdown Format for AI Agents - Suganthan Mohanadasan (GEO practitioner build + free OKF generator) - `suganthan.com/blog/open-knowledge-format` - Open Knowledge Format vs. RAG: Rethinking how AI agents get their context - HPE Community (the rederive-versus-read framing) - `community.hpe.com/t5/ai-unlocked/open-knowledge-format-vs-rag-rethinking-how-ai-agents-get-their/ba-p/7270244` - Google's Open Knowledge Format (OKF) vs. RAG - Alphamatch (the minority "replaces RAG" position) - `alphamatch.ai/blog/google-open-knowledge-format-okf-vs-rag-2026` --- ## Claude Opus 5: Pricing, Benchmarks & What Actually Changed > Claude Opus 5 launched July 24, 2026 at the same price as Opus 4.8. The five effort levels, the real cost per task, and what breaks when you upgrade. - Canonical: https://geotoolbox.ai/blog/claude-opus-5 - Published: 2026-07-25 · Updated: 2026-08-14 Claude Opus 5 launched on July 24, 2026 at exactly what Opus 4.8 cost the day before. That is the headline, and it is the least useful thing about the release. The number that decides your actual bill is a five-value parameter most teams will never set. The benchmark figures circulating in launch articles do not agree with each other, and the reason turns out to be that Anthropic quietly reissued the system card. And the two most-read hands-on takes reached opposite conclusions about what the model is like to work with. All of that is checkable. Most of it is not in the launch post. ## What Is Claude Opus 5? Claude Opus 5 is Anthropic's Opus-tier model, released July 24, 2026. It costs the same per token as Opus 4.8 and lands close enough to [Claude Fable 5](https://geotoolbox.ai/blog/fable-5-ban) that Anthropic describes it as coming near frontier intelligence at half the price. It is a [reasoning model](https://geotoolbox.ai/glossary/reasoning-model), and it is the default model on Claude Max and the strongest model available on Claude Pro. It is not on the free tier.
SpecificationClaude Opus 5
Release dateJuly 24, 2026
API model IDclaude-opus-5 (Bedrock: anthropic.claude-opus-5; Google Cloud: claude-opus-5)
Pricing$5 / million input tokens, $25 / million output tokens
Context window1M tokens (default and maximum)
Max output128K tokens (up to 300K via the Batch API with a beta header)
Knowledge cutoffMay 2026
Effort levelsFive: low, medium, high, xhigh, max (default high)
Fast modeRoughly 2.5x speed at $10 / $50 per million tokens
AvailabilityClaude API, Claude Pro / Max / Team / Enterprise, Claude Code, Bedrock, Google Cloud
One framing detail gets lost in the launch coverage. Opus 5 is not the top of Anthropic's lineup. The [model documentation](https://platform.claude.com/docs/en/about-claude/models/whats-new-opus-5) still routes the hardest workloads to Fable 5, which costs twice as much. Opus 5 is the model built to be used every day, not the one built to win every task. If you are new to the lineup entirely, our explainer on [Anthropic's Claude assistant](https://geotoolbox.ai/blog/what-is-claude-ai) covers how the tiers fit together. ## The Effort Setting Has Five Levels, Not Three Most launch coverage described the effort toggle as low, medium, or high. There are five. [Anthropic's documentation](https://platform.claude.com/docs/en/build-with-claude/effort) is unambiguous: "Claude Opus 5 supports all five effort levels." The parameter is `output_config.effort`, nested inside `output_config` rather than sitting at the top level of the request, and it defaults to `high`. Setting it to `high` produces exactly the same behavior as leaving it out. Here is the part that makes it more than a speed dial. Effort governs **all tokens in the response**, not just thinking. That includes text, explanations, and tool calls. Lower effort means Claude makes fewer tool calls, combines operations, skips the preamble, and confirms tersely. Higher effort means more tool calls, a stated plan before acting, and longer summaries.
LevelWhat it doesUse it for
lowMost efficient. Significant token savings, some capability reductionSimple tasks, classification, subagents, high volume
mediumBalanced, moderate token savingsAgentic work balancing speed, cost and performance
high (default)Equivalent to not setting the parameterComplex reasoning, difficult coding, agentic tasks
xhighExtended capability for long-horizon workAgentic runs over 30 minutes, token budgets in the millions
maxMaximum capability, no constraint on token spendThe deepest reasoning you can buy
Two things follow that are easy to miss. First, Anthropic's guidance is not "turn it up." The docs say to start at `high` and "use `low` and `medium` liberally as your primary control for token cost and response time wherever your evals show quality holds." That is the vendor telling you to spend less, which is unusual enough to take seriously, and the system card shows why. On FrontierBench v0.1, Opus 5's best result came at `xhigh`, not `max`, with "max scored similarly and within noise." Below that, the curve gets expensive fast: `high` scores 39% against `xhigh`'s 44.4% while using 19% fewer output tokens, and `low` scores 25% for 64% fewer tokens. The same pattern shows up on AutomationBench, and here Anthropic puts a price on it. Opus 5 at max effort scored 26.0%. At medium effort it scored 24% "at $0.89 cost per task, significantly outperforming both Claude Opus 4.8 and Fable 5 at less than half the cost." Two points of score for roughly double the spend, from Anthropic's own evaluation. Be careful how far you take that. On AutomationBench `max` did score highest, and Anthropic documents it as the highest-capability setting. The honest reading is narrower and still useful: the top of the dial is reliably the most expensive setting, and on two of Anthropic's benchmarks a lower one landed within a couple of points for a fraction of the spend. That is the whole argument for running an effort sweep on your own evaluations before you standardize on a level. Second, effort is not new to Opus 5. The same parameter works on Opus 4.5 and later, plus Sonnet 5 and Fable 5. What changed is how much the choice costs you, which is the next section. If you carried settings over from an earlier model, the docs are explicit that you should run a fresh sweep on your evals instead of reusing them, and that effort does not reliably shorten visible responses on Opus 5. If you want shorter answers, ask for shorter answers. ## Claude Opus 5 Pricing Opus 5 costs $5 per million input tokens and $25 per million output tokens. That is identical to Opus 4.8 and exactly half of Fable 5. The full rate card sits in [Anthropic's pricing documentation](https://platform.claude.com/docs/en/about-claude/pricing).
Component (per million tokens)Opus 5Fable 5Sonnet 5
Input$5$10$2
Output$25$50$10
Cache write (5 minute)$6.25$12.50$2.50
Cache write (1 hour)$10$20$4
Cache hits and refreshes$0.50$1$0.20
Batch API input$2.50$5$1
Batch API output$12.50$25$5
Fast mode input / output$10 / $50Not offeredNot offered
US-only inference (inference_geo)1.1x on all categories1.1x1.1x
The Sonnet 5 column ($2/$10) is the permanent rate: Anthropic first announced it as introductory pricing through August 31, 2026, then cancelled the planned September rise to $3 and $15, so the gap to Opus 5 holds. Three details in the Opus 5 column are worth pulling out. **There is no long-context surcharge.** The full 1M-token window bills at the standard rate. A 900,000-token request costs the same per token as a 9,000-token one. That is a real differentiator against models that apply a multiplier above a threshold. **Fast mode is a separate product decision.** It runs roughly 2.5 times faster at twice the price, as a research preview limited to the first-party Claude API. The Batch API and partner clouds do not offer it. **The minimum cacheable prompt dropped to 512 tokens**, down from 1,024 on Opus 4.8. If you are caching short system prompts that previously fell under the threshold, they now qualify. For a fuller picture of how these API rates compare to the subscription plans, see our breakdown of [Claude pricing across plans](https://geotoolbox.ai/blog/claude-pricing). ## The Token Price Is Flat. The Cost per Task Might Not Be. Same price as Opus 4.8 is the headline everyone ran. It is true per token, and it is not the whole bill. Two changes push token consumption up. Thinking is now on by default, where equivalent requests on Opus 4.8 ran without it. And Anthropic's prompting guide concedes that responses now run longer than on prior Opus models. More tokens per job at the same rate per token means a higher bill per job. [Artificial Analysis](https://artificialanalysis.ai/articles/opus-5), which runs its own independent evaluation suite, measured it. Their finding: Opus 5 at max effort "costs $2.03 on average per Intelligence Index task, below Claude Fable 5 (with fallback) at $2.75, but still above Claude Opus 4.8 (max) at $1.80 and Claude Sonnet 5 (max) at $1.53."
ConfigurationCost per Intelligence Index task
Claude Sonnet 5 (max)$1.53
Claude Opus 4.8 (max)$1.80
Claude Opus 5 (max)$2.03
Claude Fable 5 (with fallback)$2.75
Read that table twice. Against Fable 5, Opus 5 is the cheaper way to buy comparable intelligence, and Artificial Analysis puts the saving at 26%. Against its own predecessor, it costs **more per task** at max effort, at identical per-token pricing, because it spends more tokens getting there. The effort setting is where you get that back. Running the same Intelligence Index at max effort cost $3,835.51; at high effort, $1,973.77. The intelligence difference between those two runs was 61 versus 59, which moved Opus 5 from first of 190 models to fifth. Roughly double the spend for two index points, and whether that trade is worth it is entirely workload-dependent.
![Claude Opus 5 effort levels low, medium, high, xhigh and max, with Artificial Analysis measurements: high effort scores 59 on the Intelligence Index at 58.0 tokens per second and cost $1,973.77 to evaluate, while max effort scores 61 at 52.8 tokens per second and cost $3,835.51](/blog/claude-opus-5/effort-levels-cost.png)
The default is high.
The sharpest criticism runs further. In the Hacker News launch discussion, one commenter cited the Vals Index going "from $2.90 to $8.54, for 4% gain" between Opus 4.8 and Opus 5. No effort level was stated, and it is a single unreproduced third-party report, so treat it as a caution rather than a finding. Pushing the other way, several practitioners reported the opposite on real work, describing tasks that Opus 5 completed in substantially fewer tokens than Opus 4.8 needed. Both can be true, because cost per task depends on the task and the effort level. The honest summary: Opus 5 is cheaper per unit of intelligence than Fable 5, and at max effort it is measurably more expensive per task than Opus 4.8 despite the identical rate card. If you run agents at volume, our notes on why [token costs add up](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) apply directly here. ## Benchmarks: What Anthropic Reported vs What Others Measured Start with something that explains most of the confusion in the launch coverage. [Anthropic's announcement](https://www.anthropic.com/news/claude-opus-5) presents its benchmark comparisons as charts, with prose carrying only relative claims: Opus 5 "more than doubles" Opus 4.8 on FrontierBench, scores "three times as high as the next-best model" on ARC-AGI 3, lands "within 0.5% of Fable 5's peak" on CursorBench 3.2. The absolute numbers are in the [Claude Opus 5 system card](https://www.anthropic.com/claude-opus-5-system-card), a 190-plus-page PDF, in a plain four-column table that almost no launch coverage used. There is a catch. **Anthropic has served at least two revisions of that system card, and the FrontierBench row is not the same in both.** The revision published at launch gives Opus 4.8 18.7%, Fable 5 33.7% and GPT-5.6 Sol 37.5%. The revision the canonical link serves now gives 21.1%, 33.8% and 34.4%. Opus 5's own 43.3% is unchanged; all three of its comparators moved, and GPT-5.6 Sol swung 3.1 points. Both revisions also still carry the same sentence four pages later, in the FrontierBench section: Opus 4.8 "achieved 18.7%." So in the current revision, the summary table and the prose disagree with each other. That is why you will see both 18.7% and 21.1% quoted as Opus 4.8's score. Neither camp misread anything. They read different PDFs. Here is the table as the current revision gives it, alongside what outside parties measured independently.
Benchmark (system card)Opus 5Opus 4.8Fable 5GPT-5.6 Sol
SWE-bench Pro79.269.28064.6
SWE-bench Multimodal59.438.454.1-
DeepSWE v1.168.859.069.772.7
FrontierCode 1.1 (Main)53.446.553.547.5
FrontierBench v0.1 (revision-dependent)43.321.1 (18.7 at launch)33.8 (33.7)34.4 (37.5)
The FrontierBench row is the summary-table figure. The system card's own FrontierBench section reports 44.4% for Opus 5 at xhigh effort, noting max "scored similarly and within noise, landing at 43%."
BrowseComp90.884.387.490.4
Humanity's Last Exam (with tools)64.757.963.9-
OSWorld 2.070.655.766.162.6
HealthBench Professional59.857.4Not published (66.0 is Mythos 5)60.5
GDPval-AA v21861159317471736
AutomationBench26.017.017.418.1
Independent measurement adds a different set of numbers:
MeasurementResultMeasured by
AA Intelligence Index (max effort)61, ranked #1 of 190 modelsArtificial Analysis
AA output speed (max effort)52.8 tokens/sec, ranked #115 of 190Artificial Analysis
AA-Omniscience hallucination rate+14 points to 50% vs Opus 4.8Artificial Analysis
Overall (provisional)#1 of 215 models (85.88)BenchLM
Knowledge / agentic / multimodal#1 of 53 · #3 of 129 · #3 of 32BenchLM
Coding category#9 of 130 models (68.8)BenchLM
Three observations that the launch post does not surface. **It is the smartest model measured, and it is slow.** Artificial Analysis ranks Opus 5 first of 190 on intelligence and 115th of 190 on output speed. That is the tradeoff, and it explains why a paid fast mode exists at all. **"Best at coding" deserves an asterisk.** The aggregator BenchLM ranks Opus 5 #1 of 215 models overall, at 85.88, and its category breakdown complicates the coding pitch specifically: #1 of 53 on knowledge, #3 of 129 on agentic, #3 of 32 on multimodal, but #9 of 130 on coding. Those are all strong placings, and coding is the weakest of them rather than the strongest. Hold all of it loosely, though, and by the same standard applied to Anthropic above: BenchLM has published 79 of the 369 benchmarks it tracks for this model, the categories above are tagged "mixed sources," and it has not yet promoted Opus 5 to its verified leaderboard. **Anthropic reports its own second places. The announcement does not.** The system card is candid in a way the launch post is not. Fable 5 leads SWE-bench Pro (80 to 79.2), GPT-5.6 Sol leads DeepSWE v1.1 (72.7 to 68.8) and HealthBench Professional (60.5 to 59.8), and FrontierCode is effectively a tie at 53.5 to 53.4 in Fable's favour. That is at least four rows where Opus 5 is not first, inside a launch headlined on agentic coding, and you only find them if you open the PDF. Treat the near-ties as ties, though: the system card configures effort per benchmark, so a tenth of a point is not a ranking. ## What Actually Improved Over Opus 4.8 Strip out the benchmark noise and five changes show up repeatedly in both vendor material and practitioner reports. **Self-verification.** Opus 5 checks its own work without being asked. Anthropic describes it opening rendered pages at desktop and mobile widths, spotting layout bugs, and fixing them before handing anything back. Several developers reported the same pattern independently. **Judgment.** It pushes back. Given an unsound plan it will argue the point and propose an alternative rather than complying. Whether that reads as a feature depends heavily on how you work, which is a theme we return to below. **Knowledge freshness.** The May 2026 cutoff is four months newer than the January 2026 cutoff on both Opus 4.8 and Fable 5. For a model answering questions about its own ecosystem, that gap is larger than it sounds. **Science.** Anthropic reports gains on every life-sciences evaluation it ran, including 10.2 percentage points over Opus 4.8 on inferring molecular structures from spectroscopy data and 7.7 points on predicting how sequence variations affect protein function. These are stated as deltas, without absolute scores. **Visual output.** Anthropic showcased a wind tunnel simulation and an interactive cell illustration built entirely by the model, and customers building in this space reported the strongest animation and 3D work they had seen from an Opus model. The story doing the most work in the launch post is an anecdote rather than a benchmark. On a FrontierBench task, the model was asked to rebuild a machine part as a 3D FreeCAD model but deliberately given no way to view the drawing. It wrote its own computer vision pipeline to extract the geometry from raw pixels, then built the part. That anecdote describes the release better than any benchmark does, and it is also the most contested thing about it. Simon Willison, careful to note he had not yet tested the model, called the behavior [relentlessly proactive](https://simonwillison.net/2026/Jul/24/introducing-claude-opus-5/) from the same passage. For background on the mechanics underneath, see [how Claude works](https://geotoolbox.ai/blog/how-does-claude-work). One improvement went the other way, and it is the least comfortable number in the launch. Anthropic's system card states that the model "hallucinates factual claims slightly more than Opus 4.8, despite being more accurate overall." Artificial Analysis measured the same trade: on its AA-Omniscience benchmark, Opus 5 improves factual accuracy by 7 points over Opus 4.8 but "answers more often when uncertain," with its hallucination rate rising 14 points to 50%. More knowledge and more confident wrong answers shipped together. If you are pointing this model at research or knowledge work, that is the caveat that matters most, and it is the one the launch tables do not carry. ## The Safeguards Loosened, and That Is the Enterprise Story The change most likely to unblock an actual purchase decision is not on the benchmark chart. Opus 5's cybersecurity classifiers are less restrictive than Fable 5's. Anthropic's [announcement](https://www.anthropic.com/news/claude-opus-5) states it directly: "Based on our testing, we expect the classifiers to intervene around 85% less often than they do for Fable 5." That is an expectation from internal testing rather than a measured production figure, which is worth holding in mind, but it is Anthropic's own published number. [TechCrunch's write-up](https://techcrunch.com/2026/07/24/anthropic-launches-opus-5/) reports the same expectation. The boundary itself is clearer. Opus 5 may search for vulnerabilities in source code, and is blocked from scanning a software binary for them. Anthropic frames the capability honestly: the model "comes close to Mythos 5 at finding cybersecurity vulnerabilities" but "remains substantially behind Mythos 5 on the exploitation of those vulnerabilities." It says it avoided training Opus 5 on cyber tasks entirely, and that the gains came from general capability instead. Security researchers doing legitimate work have a further route: members of Anthropic's Cyber Verification Program get access to a version with fewer restrictions. Then the part enterprise buyers actually care about. Opus 5 is not subject to the 30-day data retention policy that covers Fable and Mythos. For any organization whose security review stalled at Fable 5's retention requirement, that is what changes the answer, and it is nowhere on the benchmark chart. ## What Breaks When You Upgrade Swapping the model string is not sufficient. Four changes will bite production code. **Thinking is on by default.** On Opus 4.8, a request without a `thinking` field ran without thinking. On Opus 5, the same request thinks. Your token consumption changes without any code change. **You cannot turn thinking off at the top two effort levels.** Setting `thinking: {"type": "disabled"}` at `xhigh` or `max` returns a 400 error. This is documented, and it is the kind of thing that surfaces in production rather than in testing. **`max_tokens` now bounds thinking plus response text together.** A limit tuned for Opus 4.8's output alone will truncate differently here. Anthropic suggests starting around 64K for `xhigh` and `max` work and tuning from there. **Refusals are not errors.** When a safety classifier declines a request, the docs are explicit that "safety classifiers return this stop reason as a normal HTTP 200 response, not an error," with `stop_reason` set to `refusal` and a `stop_details` object naming the policy category. Error handling that only watches for non-200 status codes will sail straight past it. Anthropic notes that a refused request on Opus 5 or Fable 5 can usually be served by retrying on another Claude model. There is a workflow-level version of this too. The team at Every, testing Opus 5 for a week before launch, found it "argued with instructions, stopped before the work was finished, and generally didn't play well with our existing skills and plugins." Their fix was to delete the accumulated scaffolding and start over, after which results improved sharply. Anthropic's own documentation points the same direction, advising a fresh effort sweep rather than carrying settings across. **You may not get the 1M context you read about.** Day-one users reported seeing roughly 200K through Claude Code, and the explanation is in the Claude Code docs rather than the launch post. On Max, Team and Enterprise plans, Opus is "automatically upgraded to 1M context with no additional configuration." On Pro it is not: extended context there "requires usage credits." You can force it with the `opus[1m]` alias or `/model opus[1m]`, and an admin can switch it off entirely with `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`, which strips the 1M variants from the picker. The window itself carries no price premium beyond 200K. So the 1M figure is real, and whether you have it depends on your plan and your model string. If long-context behavior matters, check which variant you are actually on before committing, and see our guide to [context window limits](https://geotoolbox.ai/blog/claude-code-context-window) for why the advertised [context window](https://geotoolbox.ai/glossary/context-window) and the usable one so often diverge. Changing the effort value between requests invalidates prompt caching, because effort shapes the rendered prompt. Pick a level at the start of a cached session and hold it. ## Opus 5 vs Fable 5 vs Sonnet 5: Which One to Use
Claude Opus 5Claude Fable 5Claude Sonnet 5
Price (in / out)$5 / $25$10 / $50$2 / $10 (permanent rate)
PositionEveryday premiumDocumented capability ceilingVolume workhorse
Knowledge cutoffMay 2026January 2026January 2026
Data retentionNo 30-day requirement30-day requirementNo 30-day requirement
Cost per task at max effort (AA)$2.03$2.75 (with fallback)$1.53
Best forComplex coding, agents, enterprise work, scienceThe hardest autonomous workEveryday coding and research at volume
So, is Opus 5 better than Fable 5? For most work, yes, at half the token price, and it outscores Fable 5 on most rows of Anthropic's own table. But the system card opens by stating the opposite in as many words: "Claude Opus 5 is not more capable overall than our most capable general-access model, Claude Fable 5." Fable still leads SWE-bench Pro, and the docs still route the highest-capability workloads to it. Anthropic does not reconcile the two for you. The practical reading: default to Opus 5, and escalate to Fable 5 only when Opus 5 measurably fails on your workload and the cost of that failure exceeds the premium. If you have never hit a wall with Opus 5, the extra $5 and $25 buys you nothing. Against [Claude Sonnet 5](https://geotoolbox.ai/blog/claude-sonnet-5) the decision is more common and cuts the other way. Sonnet 5 is the sensible default for high-volume everyday work, and Opus 5 earns its premium when accuracy on long, multi-step tasks is what you are paying for. If you are on Opus 4.8 today, the per-token rate is unchanged, but the capability gain is not automatically free: the same task can consume more thinking and response tokens, which is the unit you actually pay for. Whether that trade is worth it is the one thing no benchmark table can tell you. ## The Reviewers Cannot Agree on Its Personality Here is the strangest thing about this launch. Two teams with early access reached opposite conclusions about what Opus 5 is like to work with.
How I AI (Claire Vo)Every
Headline verdict"Brilliant (but annoying)""A hard model to love"
Personality readTimid and over-apologetic, in her companion video "I hate Opus 5. It's the best model, anyway."Pushy, opinionated, argumentative
Failure mode observedAsks for confirmation on tasks it could simply doStops early, overrides instructions, breaks existing skills
Named the problem"Claude slop" (verbosity)"A poor man's Fable"
Both are first-hand and both are early: [Claire Vo's is a podcast review](https://www.lennysnewsletter.com/p/claude-opus-5-review-this-model-is), and [Dan Shipper's is a day-one thread](https://x.com/danshipper/status/2080700057892815114) from a team with a week of pre-launch access, which he labels a "Day 0 vibe check." The overlap matters more than the contradiction. In both cases the model does something other than what it was told, and in both cases the reviewer noticed the personality before the capability. Anthropic's [prompting guide for Opus 5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5) concedes the verbosity half, noting that its "default user-facing responses run longer than prior Opus models" and advising you to prompt for length rather than expecting a lower effort setting to fix it. Which review is right matters less than what the disagreement implies: personality is now a deployment variable. It decides whether your existing prompts still work, and no benchmark score will warn you about it. ## Without Retrieval, Every Engine We Tested Failed the Question We ran a panel test on July 25, 2026, the day after launch, asking eight engine configurations the same question about Claude Opus 5.
Engine configurationRetrievalResult
GPT-5.6 Terra (OpenAI API)OnCorrect
Gemini 3.1 ProOnCorrect
Perplexity (Sonar Pro)OnCorrect
Claude (Opus 4.8)OnCorrect
ChatGPT, live web interfaceOnCorrect on date and pricing, wrong on context window
GPT-5.6 Terra (OpenAI API)OffFailed, could not verify any of it
Claude (Opus 4.8)OffFailed, denied its own existence
Gemini 3.1 ProOffFailed, declared the Claude 5 generation fictional
With web retrieval enabled, five out of five got the July 24 release date and both price points right, within roughly twenty-four hours of the announcement. Retrieval is not a guarantee, though, and our own panel produced the counter-example. One of those five stated the context window as 200,000 tokens, twice, in a bolded table, citing Reuters, The Verge and Axios for a figure none of them carried. The other four said 1M. Retrieval solved the recency problem. It did not solve the confidence problem. With retrieval disabled, three out of three failed. Not "were a bit behind." Failed. Two refused cleanly, saying they could not verify any of it. The third did something more instructive. Gemini 3.1 Pro, asked the same question in the same session minutes apart with one flag flipped, declared that Claude Opus 5, Opus 4.8 and Fable 5 all did not exist, stated that Anthropic's most recent generation was the Claude 3 and 3.5 family, asserted that Anthropic models have no effort parameter, then supplied Claude 3 Opus specifications from March 2024 as the current answer. It attributed the premise of the question to "fictional internet speculation." Claude Opus 4.8, with retrieval off, denied its own existence. On the same day, `claude-opus-5` was still absent from DataForSEO's LLM model catalog, which is the sort of infrastructure lag that decides whether a tool can even see a new model. None of this is a knock on the engines. It is the mechanic underneath every AI answer about anything recent: the response is a **retrieval result, not a memory**. Same weights, same prompt, same session. The only variable was whether the model could go and look, and that variable produced the difference between a correct answer and a confidently [hallucinated](https://geotoolbox.ai/glossary/ai-hallucination) one. There is one more result from the same test, and it cuts against the easy conclusion. Most of what the engines actually cited was not Anthropic's documentation. Across the panel, first-party citations were rare: one engine returned 24 sources with a single anthropic.com link, and that link pointed at Opus 4.5. The rest were third-party pages, several of them written before the model shipped. Anthropic's own docs are reachable, crawlable and correct, and the engines largely went elsewhere anyway. So being crawlable is the floor, not the finish line. It decides whether you are eligible to be cited, not whether you will be. When an engine answers a question about your product, your category, or your pricing, it is reading something, and if your pages are not reachable by the crawlers that feed those answers you are not in the running at all. That gap between "published" and "retrievable" is the problem geotoolbox exists to find. ## The Setting to Change First If you take one thing from this, make it the effort parameter. The price did not change, the capability went up, and the variable that decides what you actually pay is a string most teams will never set. Anthropic's advice is to start at `high` and reach for `low` and `medium` liberally wherever your evals hold. Independent evaluation runs cost roughly twice as much at `max` as at `high` for a two-point intelligence difference, and on Anthropic's FrontierBench numbers `max` is not even the peak. Running everything at `max` because it sounds better is the most expensive mistake available here. So the next question is not really "is it better." The token rate alone favours testing the upgrade, but it does not settle the deployment decision: cost per completed task, refusal behaviour, the hallucination trade and the migration work all sit on the other side of it. The question worth answering first is which effort level your workload actually needs, and that means running your own evals across at least three of the five levels before you standardize on one. While you are auditing what the model can do, it is worth checking what the models can see. If you also own the pages your product gets documented on, and AI engines answer questions about your business from whatever they can retrieve, the first thing to confirm is that they can reach your pages at all. The free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) we built at geotoolbox tests whether the crawlers behind ChatGPT, Claude, Perplexity and Google's AI surfaces are actually allowed to fetch your site. It is a faster answer than most people expect, and occasionally an unwelcome one. ## Frequently Asked Questions ### Is Claude Opus 5 available on the free plan? No. Claude Opus 5 requires a paid plan. It is the default model on Claude Max and the strongest model available on Claude Pro, and it is also available on Team, Enterprise, Claude Code and the API. Free users get Claude Sonnet 5 instead. ### How much does Claude Opus 5 cost? $5 per million input tokens and $25 per million output tokens, which is identical to Opus 4.8 and half of Claude Fable 5. Batch API pricing is $2.50 and $12.50, cache hits are $0.50, and fast mode doubles the base rate to $10 and $50 for roughly 2.5 times the speed. There is no surcharge for using the full 1M-token context window. ### Is Claude Opus 5 better than Claude Fable 5? On most of Anthropic's own published benchmarks, yes, at half the token price. But Anthropic still directs workloads needing the highest available capability to Fable 5, and states that Opus 5 is not more capable overall. The practical approach is to default to Opus 5 and escalate to Fable 5 only when Opus 5 measurably fails on your specific workload. ### What is the default effort level for Claude Opus 5? The API default is `high`, and setting `high` explicitly produces exactly the same behavior as omitting the parameter. Opus 5 supports all five levels: `low`, `medium`, `high`, `xhigh` and `max`. You pass it as `output_config.effort`. ### Does Claude Opus 5 work in Claude Code? Yes, and `high` is the default there as well. The context window depends on your plan: Max, Team and Enterprise get the 1M window automatically, while Pro requires usage credits for it. You can select it explicitly with `/model opus[1m]`, and it carries no price premium beyond 200K. ### Is Claude Opus 5 better than GPT-5.6? On Anthropic's published comparisons, Opus 5 leads GPT-5.6 Sol on most rows: FrontierBench (43.3 to 34.4 in the current system-card revision), OSWorld 2.0 (70.6 to 62.6), SWE-bench Pro (79.2 to 64.6), GDPval-AA v2 (1861 to 1736) and AutomationBench (26.0 to 18.1). GPT-5.6 Sol leads on DeepSWE v1.1 (72.7 to 68.8) and HealthBench Professional (60.5 to 59.8), and effectively ties BrowseComp. These are Anthropic's numbers on Anthropic's harness, so treat them as directional and run your own evaluation on the work you actually do. ### How much does a Claude Opus subscription cost? Opus 5 is not sold separately. It comes with Claude Pro at $20 a month, where it is the strongest model available, and with Claude Max at $100 or $200 a month, where it is the default. Team and Enterprise plans include it too. Paying by the token through the API instead costs $5 per million input and $25 per million output, on a completely separate bill from the subscription. ### Is Claude Opus 5 worth upgrading to from Opus 4.8? The per-token price is unchanged, but the capability gain is not automatically free: thinking is now on by default and responses run longer, so the same task can consume more tokens. Test on your own workload at `medium` and `high` before assuming the bill stays flat, and budget time for the breaking changes around `thinking` and `max_tokens`. ## Sources - Introducing Claude Opus 5 - Anthropic, July 24, 2026 - `anthropic.com/news/claude-opus-5` - Claude Opus 5 System Card - Anthropic, July 24, 2026 (two revisions compared: `c5fbac3f`, 193pp, and the canonical `b514064a`, 194pp) - `anthropic.com/claude-opus-5-system-card` - Effort - Claude Platform Docs, Anthropic - `platform.claude.com/docs/en/build-with-claude/effort` - What's new in Claude Opus 5 - Claude Platform Docs, Anthropic - `platform.claude.com/docs/en/about-claude/models/whats-new-opus-5` - Pricing - Claude Platform Docs, Anthropic - `platform.claude.com/docs/en/about-claude/pricing` - Prompting Claude Opus 5 - Claude Platform Docs, Anthropic - `platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5` - Opus 5: Fable 5 level intelligence at a lower cost per task - Artificial Analysis, July 24, 2026 - `artificialanalysis.ai/articles/opus-5` - Claude Opus 5 (max) and (high) model profiles - Artificial Analysis, July 2026 - `artificialanalysis.ai/models/claude-opus-5` - Claude Opus 5 benchmark profile - BenchLM, July 2026 - `benchlm.ai/models/claude-opus-5` - Claude Opus 5 review: this model is brilliant (but annoying) - Claire Vo, Lenny's Newsletter, July 24, 2026 - `lennysnewsletter.com/p/claude-opus-5-review-this-model-is` - Day-one thread on testing Claude Opus 5 - Dan Shipper, Every, July 24, 2026 - `x.com/danshipper/status/2080700057892815114` - Introducing Claude Opus 5 - Simon Willison's Weblog, July 24, 2026 - `simonwillison.net/2026/Jul/24/introducing-claude-opus-5/` - Anthropic launches Opus 5 - Russell Brandom, TechCrunch, July 24, 2026 - `techcrunch.com/2026/07/24/anthropic-launches-opus-5/` - Introducing Claude Opus 5 (launch discussion) - Hacker News, July 24, 2026 - `news.ycombinator.com/item?id=49038433` --- ## The Best AI Search Engines in August 2026 > The best AI search engines in August 2026, ranked by what you actually use them for, with vendor-checked pricing and an honest look at citations. - Canonical: https://geotoolbox.ai/blog/best-ai-search-engines - Published: 2026-07-23 · Updated: 2026-08-13 Most "best AI search engine" lists rank on vibes and ship prices that were stale by the time they published. This one sorts the picks by what you actually use each tool for, and every price was checked against the vendor's current pricing in July 2026. It is also honest about the flaw every AI-powered search engine still shares: citations you cannot fully trust. Here are the ones worth your time, what each is genuinely best at, and where to stay skeptical. ## The Best AI Search Engines at a Glance For most people who want answers with sources they can check, Perplexity is the one to start with. That is not a unanimous verdict, and it should not be. PCMag's testers rank Google's AI Mode first; Zapier and others put Perplexity on top. The honest answer is that the best AI search engine depends on what you are doing with it, so the table below sorts by use case rather than crowning a single winner.
EngineBest forFree tierPaid (from)Shows citationsReal-time web
PerplexityCited researchYesPro $20/moYesYes
ChatGPT SearchAll-around assistantYesGo $8/mo, Plus $20/moYesYes
Google AI Mode & GeminiEveryday, local, shoppingYes (in Search)Gemini AI Plus $4.99/moPartialYes
Microsoft CopilotMicrosoft 365 usersYesBundled in 365 Premium $19.99/moYesYes
ClaudeWriting and analysisYesPro $20/moYes (web on)Yes
GrokReal-time X and socialYesSuperGrok $30/moYesYes
Brave SearchPrivacy, freeYesPremium $3/moYesYes
KagiPaid, ad-free power useTrial onlyStarter $5/moYesYes
You.comMulti-model workspaceYesPro ~$20/moYesYes
Prices are US monthly rates checked against each vendor's current pricing on July 23, 2026, before tax. The "Paid (from)" column shows the cheapest paid tier. AI search pricing changes often, so treat it as a starting point and confirm before you subscribe. "Shows citations" means the tool displays sources, not that they are always accurate, which is a distinction the citations section below gets into. ## How We Picked and Priced These Five things separate a useful AI search engine from a demo that falls apart on real work. **Citation quality and verifiability.** Not whether it shows sources, but whether those sources actually support the claim. This is where most tools quietly fail, and it gets its own section below. **Use-case fit.** A tool built for cited research behaves differently from one built for coding or for shopping. Ranking them on one axis hides that, so the picks are organized by what you are trying to do. **Real-time access.** Whether the tool retrieves live pages or answers from stale training data. Most now retrieve, but the quality of that retrieval varies more than the model name suggests. **Free-versus-paid value.** What the free tier actually gives you before the limits bite, and whether the paid tier earns its price. **Privacy.** What the tool logs, whether it trains on your queries, and whether you can turn that off. On pricing, every number here was checked against the provider's current page in July 2026. That is not busywork: Microsoft retired standalone Copilot Pro in late 2025, yet lists still quote its old $20 price. A guide that ranks its own product as the best "search engine" is running an ad, not a review, so those did not make the cut either. ## The AI Search Engines, Reviewed Each entry names what the tool is best at, the one thing to watch for, and its verified price. Links go to a full breakdown where we have one. ### Perplexity: Best for Cited Research Perplexity was built around the thing other tools bolt on later: answers come with numbered inline citations, and there is a dedicated sources view. Ask it a messy research question and it returns a structured answer with the receipts attached, which is why it is the default pick for anyone who needs to verify what they read. Collections let you keep ongoing research organized by project. The deeper walkthrough is in our guide to [what Perplexity is](https://geotoolbox.ai/blog/what-is-perplexity). Watch for two things. Perplexity tightened its Pro and Deep Research quotas in 2026, and heavy users noticed. It has also drawn sustained criticism over how its crawler pulls publisher content, an ethics question Zapier flags that most reviews skip. We still rank it first on citation quality, but if publisher ethics is your deciding factor, Brave and Kagi have drawn far less of that criticism. The free tier is generous for casual use; [Perplexity's pricing](https://geotoolbox.ai/blog/perplexity-pricing) runs $20/mo for Pro and $200/mo for Max. ### ChatGPT Search: Best All-Around Assistant ChatGPT Search is the one most people already have open. It handles conversational, multi-step questions well, holds context across follow-ups, and can pull in your own uploaded documents alongside live web results. If you want one tool that searches, drafts, and reasons in the same thread, this is it. The mechanism behind it, and why the model itself never actually browses, is worth understanding: see [how ChatGPT Search works](https://geotoolbox.ai/blog/how-chatgpt-search-works). And if you are weighing assistants as assistants rather than search engines, our [best ChatGPT alternatives](https://geotoolbox.ai/blog/chatgpt-alternatives) roundup sorts that field by job. Its weak spot is the same one the whole category shares. Citation accuracy varies, and it will confidently answer questions it should decline. The free tier now includes web search; paid [ChatGPT pricing](https://geotoolbox.ai/blog/chatgpt-pricing) starts at Go $8/mo, then Plus $20/mo, and Pro from $100/mo (up to $200/mo for the highest limits). ### Google AI Mode, AI Overviews and Gemini: Best for Everyday, Local, and the Google Ecosystem Google is where AI search meets the search index nobody else has. [AI Mode](https://geotoolbox.ai/blog/what-is-google-ai-mode) gives you a conversational answer inside a normal Google search, [AI Overviews](https://geotoolbox.ai/blog/what-are-google-ai-overviews) summarize at the top of the results, and [Gemini](https://geotoolbox.ai/blog/what-is-gemini) is the standalone assistant with the deepest ties to Maps, Shopping, and your Google account. For everyday questions, local intent, and price comparison, nothing matches its coverage. The trade-off is trust, which is why the table rates Google only "Partial" on citations: AI Overviews attribute sources weakly, and Google's confidently-wrong misfires are well documented. AI Mode and AI Overviews are free inside Google search; only the standalone Gemini app charges, from AI Plus at $4.99/mo to AI Pro at $19.99/mo and up to Ultra. ### Microsoft Copilot: Best for Microsoft 365 Users Copilot blends AI answers with Bing's live results and, more to the point, sits inside Word, Excel, Outlook, and Teams. If your work already lives in Microsoft 365, that integration is the reason to use it. The full picture is in our guide to [what Copilot is](https://geotoolbox.ai/blog/what-is-copilot). As a pure search tool it is competent rather than exceptional. Answer depth varies by topic, the interface is busier than Perplexity's, and the best features are the ones wired into Office rather than the open-web search itself. Here is the pricing catch worth knowing. Microsoft retired the standalone $20/mo Copilot Pro in late 2025; support for existing subscribers ended on August 1, 2026. Its consumer features now come bundled in Microsoft 365 Premium at $19.99/mo, and the business tier is $30 per user per month. The free consumer Copilot still handles Bing-grounded web search. See [Copilot pricing](https://geotoolbox.ai/blog/copilot-pricing) for the current tiers. ### Claude: Best for Writing and Analysis Claude is less a search engine than a reasoning tool that can search. With web search on, it pulls live sources, but its real strength is what it does with them: careful analysis, long-document work, and writing that needs judgment rather than just retrieval. For synthesizing a dense topic into something coherent, it is hard to beat. More on the assistant in our guide to [what Claude is](https://geotoolbox.ai/blog/what-is-claude-ai). The thing to remember is that web search has to be enabled to get current results; with it off, you are talking to training data. The free tier includes web search; [Claude pricing](https://geotoolbox.ai/blog/claude-pricing) is $20/mo for Pro and from $100/mo for Max. ### Grok: Best for Real-Time X and Social Grok's edge is native, real-time access to the pulse of X, which no other tool has as directly. For breaking events, social sentiment, and what people are saying right now, it surfaces things the others miss. Our guide covers [what Grok is](https://geotoolbox.ai/blog/what-is-grok) in full. The caution here is real and specific. In the March 2025 Tow Center citation study, the then-current Grok-3 returned incorrect source attributions 94% of the time, the worst of any engine tested. xAI has since shipped Grok 4.5 and 4.6, and nobody has independently re-tested them on citations since, so treat real-time Grok as fast but unproven on citations, and verify what matters. Free access is rate-limited; [Grok pricing](https://geotoolbox.ai/blog/grok-pricing) runs SuperGrok at $30/mo and X Premium+ at $40/mo. ### Brave Search: Best Privacy-First Free Option Brave runs its own independent search index, does not track you, and layers AI answers on top of results you can still click through to. If you want AI search without the data collection, it is the cleanest free option. It is also fully usable without an account. The trade-off is index depth. Brave's independent index is smaller than Google's, so it can feel thinner on obscure long-tail queries. Search is free; Premium removes ads for $3/mo. ### Kagi: Best Paid, Ad-Free Search Kagi is the anti-spam option. You pay, so it has no incentive to serve ads or optimize for engagement, and the results feel notably less polluted by low-quality SEO content. You can raise or bury specific domains, pick your AI assistant, and use lenses to scope searches. Because it is funded by subscriptions rather than ads, it has no ad business to feed and does not build an advertising profile from your searches. For people tired of fighting their search engine, it is a relief. The obvious catch is that it is paid-only. There is a limited free trial, then Starter is $5/mo for 300 searches, Professional is $10/mo for unlimited, and Ultimate is $25/mo with premium AI. That last number matters because some lists still quote a much higher figure. ### You.com: Best Multi-Model Workspace You.com lets you switch between models and modes in one place, blending chat, search, and productivity tools like drafting and code. For creators who turn web research into output without leaving the tab, that flexibility is the draw. The downsides are a smaller index than Google and enough modes that new users can feel lost. There is a free tier; Pro runs around $20/mo. ### Also Worth Knowing A few narrower picks fill specific gaps. **DeepSeek** offers strong free reasoning, but its data is processed in China, which is a non-starter for many companies. **Consensus** searches only peer-reviewed papers and shows where scientific agreement actually lies, making it the pick for academic work. **Phind** is tuned for developers and returns working code over prose. **Mistral's Vibe** (formerly Le Chat) is a capable European option with a privacy-friendly stance. **Komo** and **Andi** both offer clean, ad-free, no-signup free search, and for privacy specifically, **DuckDuckGo's** assistant keeps prompts out of a tracking profile. The newest shape of the category is the AI browser, where the answer engine is built into the browser itself rather than a separate site. Perplexity's [Comet](https://geotoolbox.ai/blog/perplexity-comet) and ChatGPT's Atlas are the ones to watch as consumer AI search keeps moving in that direction. ## The One Thing Every AI Search Engine Still Gets Wrong: Citations A citation is not proof. This is the single most important thing to understand before you trust any of these tools, and it is the complaint that comes up most often from people who use them daily: a confident answer, a tidy source link, and a source that does not actually say what the answer claims. The best data on this comes from the [Tow Center for Digital Journalism](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php). Researchers fed eight AI search tools direct excerpts from news articles and asked each to identify the original source. Across 1,600 queries, the tools got the citation wrong more than 60% of the time.
EngineIncorrect citationsTest
Perplexity (best performer)37%Tow Center, 2025
ChatGPT Search67%Tow Center, 2025
Grok-3 Search (worst performer)94%Tow Center, 2025
All 8 tools (average)Over 60%Tow Center, 2025
Those are the model versions tested in early 2025, and the engines have all updated since, with no comparable independent re-test published. Read the table as evidence that the problem is real and widespread, not as today's exact scoreboard. An earlier 2024 Tow Center study of ChatGPT alone found it misattributed 153 of 200 quotes, so this is not a one-off. The counterintuitive finding: the paid tiers were worse on this measure. Perplexity Pro and Grok-3 answered more prompts, but they did it by giving definitive, confidently wrong answers instead of declining, where the free versions were more likely to admit they did not know. Paying more bought more confidence and fewer "I don't know" responses, not more reliable citations. None of this means AI search is useless. It means the citation is a starting point, not a verdict. Verifying one takes about twenty seconds: open the linked source, use Ctrl-F to find the exact claim, and check the date. Be most skeptical of clean round-number statistics and anything with real stakes. If you want the mechanism behind why these tools cite the way they do, see [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work). ## It's the Retrieval, Not the Model Badge Ask the same question in two tools running the same underlying model and you can get two different answers. That surprises people, because the marketing is all about which model powers what. In practice, the model matters less than the retrieval. Here is why. An AI search engine does not think its way to an answer from memory. It runs your query through a search step, pulls back a set of web pages, and feeds that text to the language model, which then writes an answer from what it was handed. This pattern is called [retrieval-augmented generation, or RAG](https://geotoolbox.ai/blog/what-is-rag). For anything current, the model can only ground its answer in what the retrieval step gave it. If the tool fetches the wrong pages, or thin snippets, or too few of them, a great model still produces a weak answer. Good tools also break your one question into several sub-questions and search each, a technique called [query fan-out](https://geotoolbox.ai/blog/query-fan-out). It is a big part of why Perplexity and Google feel more thorough than a raw chatbot on the same model. The flip side is that retrieval is only as good as the pages it grabs, and it will grab a joke as readily as a study. This is how Google's AI Overviews ended up telling people to put glue on pizza and eat a rock: the retrieval step pulled a Reddit joke and a satire site, and the model repeated them as fact because it cannot reliably tell sarcasm or satire from a real source. A tool with a weak retrieval layer does not just miss good sources, it confidently surfaces bad ones. The practical takeaway: judge a tool by its sources, not its model badge. When you compare two engines, look at what each one actually retrieved and cited for the same query. That tells you far more than the model name in the marketing. ## What About Deep Research Modes? Most of the major tools now ship a "Deep Research" mode: Perplexity, ChatGPT, Gemini, and Grok each have a version. Instead of one quick search, it runs many, reads dozens of pages, and returns a long report with a citation list that can run past fifty sources. For a genuine research task, it is the most useful thing these tools do. It is also where the citation problem compounds. A fifty-source report is fifty chances to cite something the source does not actually say, and no one has time to open all fifty. The efficient check is to verify the load-bearing claims first, the two or three facts the conclusion actually rests on, plus any suspiciously tidy statistic, and to skim the rest. If those hold up, treat the report as a usable starting point rather than a finished answer. If they do not, the length was hiding the problem, not solving it. ## When You Should Still Use Google AI search is built for synthesis: pulling scattered information into one answer. For a lot of searches, that is the wrong job. Classic search still wins when you want to: - Reach a specific site or login you already know by name (navigational search) - Find local results, hours, directions, or a map ("dentist near me") - Compare prices or shop, where you want the actual listings, not a summary - Get to a primary source fast without a layer of interpretation in between The data backs up keeping both. Despite years of AI hype, [SparkToro found](https://sparktoro.com/blog/new-research-20-of-americans-use-ai-tools-10x-month-but-growth-is-slowing-and-traditional-search-hasnt-dipped/) that 95% of Americans still use a search engine every month, and traditional search usage has dipped by less than a percentage point in two and a half years. Stranger still, people who start using ChatGPT tend to search Google more afterward, not less. AI search is turning out to be a supplement, not a replacement. The [broader picture of where AI search actually stands](https://geotoolbox.ai/blog/state-of-ai-search-2026) is less dramatic than either side claims. ## How to Choose the Right One for You The right AI search engine is the one that fits the task in front of you. Match the task to the tool:
![Decision map matching AI search engines to use cases: Perplexity for cited research, Google AI Mode for everyday and local, ChatGPT Search for a do-everything assistant, Copilot for Microsoft 365, Claude for writing and analysis, Grok for real-time social, You.com for turning research into drafts, Brave for free privacy, DuckDuckGo for zero tracking, Kagi for paid privacy, Consensus for academic sources, and Phind for code.](/blog/best-ai-search-engines/fig-usecase-map.png)
Common use cases and a strong pick for each.
- **Cited research you need to verify:** Perplexity - **Everyday questions, local, and shopping:** Google AI Mode - **A conversational, do-everything assistant:** ChatGPT Search - **Work inside Microsoft 365:** Copilot - **Writing and careful analysis:** Claude - **Privacy, free:** Brave. Zero tracking: DuckDuckGo. Paid and spam-free: Kagi - **Real-time social and breaking events:** Grok - **Turning web research into drafts:** You.com - **Academic and peer-reviewed sources:** Consensus - **Code and developer questions:** Phind Most people who use these seriously end up with two or three, not one: a research tool, a general assistant, and Google for everything navigational. Start free, find where the limits bite, and pay only for the tier you actually run into. Once you have picked the engines you rely on, there is a second question worth asking, especially if you run a business or a website: when someone asks one of these tools about your industry, does your brand show up in the answer, and is what they say about you even right? Given that these tools get citations wrong more than 60% of the time, that is not a small question. It is a different problem from choosing a search engine, and the one geotoolbox exists to solve. We are not another search engine. We measure whether the engines above actually cite you, so you can [check how visible your site is to AI search](https://geotoolbox.ai/tools/ai-readiness) and see where you are being left out. Like the tools it measures, it samples engines that give slightly different answers each time, so read it as a signal to act on rather than a fixed score. If you want to go further, our guide to [tracking brand mentions in AI search](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search) walks through the ongoing version of that check. ## Frequently Asked Questions ### What is the best AI search engine? Perplexity is the best starting point for most people who need verifiable sources, and Google AI Mode is hard to beat for everyday, local, and shopping queries. For a conversational assistant that also searches, use ChatGPT Search. There is no single winner because the tools are built for different jobs, so the right pick depends on what you are doing. ### Are AI search engines free? Most have a usable free tier, including Perplexity, ChatGPT Search, Google, Copilot, Claude, Grok, Brave, and You.com. Kagi is the main exception: it is paid-only, with a limited trial before plans start at $5/mo. For casual use, the free tiers are usually enough. Paid tiers mainly buy higher limits and faster or more advanced models. ### Which AI search engine is most accurate? Accuracy and citations are not the same thing, and no engine is reliably accurate on its own. In the Tow Center's testing, source attributions came back wrong more than 60% of the time on average, ranging from 37% for the best performer to 94% for the worst. Perplexity scored best on citations, but you should still verify anything that matters rather than trusting any single answer. ### Does AI search replace Google? Not for most searches. AI search is good at synthesizing information into one answer, but classic search still wins for navigation, local results, shopping, and reaching primary sources quickly. The data shows people are adding AI search alongside Google, not dropping Google. Using both, and knowing which to reach for, beats picking one. ### Is Perplexity Pro or ChatGPT Plus worth paying for? Worth it once you hit the free tier's limits regularly. The roughly $20/mo tier on either tool raises usage caps and adds more capable models. If you only search a few times a day, the free version is usually fine. Try free first, and upgrade only when the limits actively get in your way. ### How do I know if an AI answer's sources are real? Click through to the source and search it for the exact claim the answer made. If the page does not contain that claim, or the link is dead, treat the answer as unverified. Round-number statistics and anything high-stakes deserve the most scrutiny, and a quick check of the source's date keeps you from repeating something that is no longer true. ### Do AI search engines train on my searches? Most consumer plans do by default, but the major tools let you opt out, and business and enterprise plans usually do not train on your data at all. ChatGPT, Gemini, Claude, and Copilot each have a setting to turn off training on your data, Perplexity and Kagi offer similar controls, and Brave and DuckDuckGo are built not to keep your queries in the first place. Defaults vary by plan and region, so if you search sensitive or client information, check the setting and turn training off before you start. ## Sources - AI Search Has a Citation Problem - Tow Center for Digital Journalism, Columbia Journalism Review, March 2025 - `cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php` - How ChatGPT Search (Mis)represents Publisher Content - Tow Center for Digital Journalism, Columbia Journalism Review, 2024 - `cjr.org/tow_center/how-chatgpt-misrepresents-publisher-content.php` - New Research: 20% of Americans Use AI Tools 10X+/Month, but Growth Is Slowing and Traditional Search Hasn't Dipped - Rand Fishkin, SparkToro, August 2025 - `sparktoro.com/blog/new-research-20-of-americans-use-ai-tools-10x-month-but-growth-is-slowing-and-traditional-search-hasnt-dipped/` - Microsoft Launches 365 Premium for Consumers, Retires Copilot Pro - PCWorld, 2026 - `pcworld.com/article/2925962/microsoft-adds-microsoft-365-premium-cuts-copilot-pro.html` - The 4 Best AI Search Engines in 2026 - Zapier, July 2026 - `zapier.com/blog/best-ai-search-engine/` - The Best AI Search Engines We've Tested for 2026 - PCMag, June 2026 - `pcmag.com/picks/the-best-ai-search-engines` - Perplexity pricing - Perplexity - `perplexity.ai/pro` - ChatGPT pricing - OpenAI - `openai.com/chatgpt/pricing` - Google AI plans (Gemini AI Plus, Pro, Ultra) - Google - `one.google.com/about/google-ai-plans` - Claude pricing - Anthropic - `claude.com/pricing` - Grok / SuperGrok pricing - xAI - `x.ai` - Brave Search Premium - Brave - `brave.com/search/premium` - You.com pricing - You.com - `you.com` - Kagi pricing - Kagi - `kagi.com/pricing` --- ## The Best ChatGPT Alternatives in 2026, Sorted by Job > The best ChatGPT alternatives in 2026, sorted by what you use them for, with vendor-checked July pricing: Claude, Gemini, Perplexity, DeepSeek, and more. - Canonical: https://geotoolbox.ai/blog/chatgpt-alternatives - Published: 2026-07-23 · Updated: 2026-08-08 People leave ChatGPT for concrete reasons: the answers can feel less sharp than they used to, the $20 plan runs into rate limits mid-task, and the default settings train on what you type. But the bigger shift is quieter. ChatGPT is no longer the only place people ask questions. Many now spread their prompts across Claude, Gemini, Perplexity, and Copilot, and the "best ChatGPT alternative" depends entirely on the job in front of you. This guide sorts the alternatives by use case instead of crowning one winner. Every price was checked against the vendor's current pricing in July 2026, and each pick comes with its real tradeoff, not just the upside. ## The Best ChatGPT Alternatives at a Glance There is no single best alternative, and any list that names one is selling you something. The ranking flips depending on whether you are writing, researching, coding, or just want an answer without paying. The table below sorts by the job, with the catch attached to each pick. Prices are US monthly list rates checked against each vendor's pricing on July 23, 2026, before tax, and show the cheapest paid tier. Every tool below has a usable free tier. AI pricing moves fast, so confirm before you subscribe.
ToolBest forFree tierPaid (from)Watch for
ClaudeWriting, reasoning, long docsYes$20/mo ProTighter usage caps; cautious refusals
Google GeminiGoogle Workspace, big contextYes$4.99/mo (AI Plus)Ties you deeper into your Google account
PerplexityResearch with citationsYes$20/mo ProCited is not the same as accurate
Microsoft CopilotMicrosoft 365 usersYes$19.99/mo (in M365 Premium)Standalone Copilot Pro was retired
DeepSeekBudget reasoning, open weightsYesNone (API pay-per-use)Data is processed in China
GrokReal-time and X/TwitterYes$30/mo SuperGrokCitation reliability is unproven
Mistral's VibeEuropean, privacy-leaningYes$14.99/mo ProSmaller ecosystem than the leaders
Meta AICasual use inside apps you haveYesNoneTied to Meta's data ecosystem
Open-source, localPrivate, offline useYesHardware cost onlySetup and hardware are the real cost
If your question is specifically about AI-powered search rather than a general assistant, our guide to the best AI search engines covers that lane in more depth. ## Which ChatGPT Alternative for Which Job Start from the task, not the brand. Most people who switch well end up using two tools: a main assistant and a specialist they open for one kind of work. Match your most common job to the pick below, then read that tool's section for the tradeoff.
![Decision map matching ChatGPT alternatives to nine jobs: Claude for writing and reasoning, Perplexity for cited research, Gemini for Google users and for images, Copilot for Microsoft 365, DeepSeek for budget coding, Grok for real-time and X, DuckDuckGo AI or HuggingChat for a no-login answer, and a local model for private or offline use.](/blog/chatgpt-alternatives/fig-job-map.png)
Start from the job, then read that tool's tradeoff below.
What you are doingStart hereWhy
Long writing and careful reasoningClaudeHolds a long thread; reads long documents
Research you need to citePerplexityAnswers come with linked sources
You live in Gmail, Docs, and SheetsGeminiBuilt into Workspace, free in Search
You live in Word, Outlook, and TeamsCopilotSits inside Microsoft 365
Coding on a budgetDeepSeekStrong reasoning, free, open weights
Real-time takes and X postsGrokWired into the live X feed
Generating or editing imagesGeminiNative image handling; ChatGPT's own image tools also lead here
A quick answer with no loginDuckDuckGo AI or HuggingChatInstant, no account needed
Private or offline workA local modelRuns on your machine, nothing leaves it
## Claude: Best for Writing, Reasoning, and Long Documents If you write for a living or work through problems that need a few steps of thought, Claude is the alternative most people settle on. It writes in a more natural voice than ChatGPT out of the box, and it holds a long thread without losing the plot. It reads long documents comfortably, with a context window around 200,000 tokens on the consumer plans, so you can drop in a full contract or a large codebase and ask questions across all of it. The free tier now includes web search. Paid Claude pricing starts at $20 a month for Pro, with a Max tier from $100 a month for heavy use. The tradeoff is caution. Claude refuses more edge-case requests than its rivals, and its free usage caps are tighter, so a long session can stop you mid-task. If you want a head-to-head on where each one wins, we compared Claude versus ChatGPT in detail. ## Google Gemini: Best for Google Workspace and Multimodal Tasks If your day runs through Gmail, Docs, and Sheets, Gemini is the natural switch. It is built into Google Workspace and answers for free inside Google Search, so you are already using a version of it. It also takes one of the largest context windows of any consumer assistant and handles images, audio, and video natively, which makes it strong for pulling apart a long PDF or a screen recording. Gemini is free in Search. The standalone app starts at $4.99 a month for the AI Plus tier and runs to roughly $19.99 for the Pro tier, per Gemini pricing. The catch is that leaning on Gemini ties you deeper into your Google account, which is the opposite of what privacy-minded switchers want. Our Gemini versus ChatGPT comparison covers where it pulls ahead and where it still trails. ## Perplexity: Best for Research You Can Cite The habit ChatGPT never broke people of is pasting its answers back into Google to check them. Perplexity fixes that by design. It searches the live web and returns answers with the sources linked inline, so you can click through and verify instead of trusting a confident paragraph. For research, competitive checks, and anything you have to defend, that is the whole game. The free tier is generous for casual use. Pro runs $20 a month and Max is $200, per Perplexity pricing. Perplexity tightened its Pro and Deep Research quotas in 2026, so heavy users hit limits sooner than they used to. One honest warning that applies to every tool here: showing a citation is not the same as being right. A linked source tells you where an answer came from, not whether the tool read it correctly, and a confident tone is not evidence either way. Treat the citations as a starting point to check, not a guarantee. See ChatGPT versus Perplexity for the full breakdown. ## Microsoft Copilot: Best If You Live in Microsoft 365 If your work happens in Word, Excel, Outlook, and Teams, Copilot is the assistant that already sits inside those apps. The free consumer version handles Bing-grounded web search, and the paid version puts AI drafting and summarizing directly in the Office ribbon. Here is where most lists are already wrong. Microsoft retired the standalone $20-a-month Copilot Pro subscription in late 2025, yet plenty of "alternatives" articles still quote that dead price. According to [PCWorld](https://www.pcworld.com/article/2925962/microsoft-adds-microsoft-365-premium-cuts-copilot-pro.html), Copilot Pro's features were folded into a new consumer plan, Microsoft 365 Premium, at $19.99 a month. [Microsoft's own support page](https://support.microsoft.com/en-us/microsoft-365-copilot/about-microsoft-copilot-pro) set an August 1, 2026 support-end date for existing Copilot Pro subscribers; after that date the plan is gone for good. For teams, the [Microsoft 365 Copilot](https://www.microsoft.com/en-us/microsoft-365-copilot/pricing) business add-on runs from about $18 to $30 per user per month depending on the plan. Copilot only earns its price if you already live in the Microsoft stack; outside it, there is less for the assistant to grab onto. Our Microsoft Copilot versus ChatGPT piece weighs it against the original. ## DeepSeek: Best for Budget Reasoning and Open Weights DeepSeek is the pick when you want frontier-grade reasoning without a subscription. Its chat is free, its reasoning and coding are genuinely competitive, and its API is a fraction of the cost of the closed frontier models, which is why developers reach for it on high-volume work. Because its weights are open, you can also run it yourself instead of renting access, which no closed model allows. The real caveat is trust. DeepSeek's data is processed in China, and several governments, including Italy, Australia, Taiwan, and South Korea, have restricted or banned it in official use over exactly that concern. For a company handling regulated or sensitive data, that alone rules it out, no matter how good the model is. For the cost details, see DeepSeek pricing. If the open-weights angle is what draws you, we explain open weights versus open source and why the difference matters. ## Grok: Best for Real-Time and X Grok, from Elon Musk's xAI, is wired directly into the live X feed, so it is the assistant to ask about a story that broke an hour ago or the current mood on a topic. That real-time hookup is its one clear edge over ChatGPT, which sees the web through a slower search layer. Free access is rate-limited. Paid Grok pricing runs SuperGrok at $30 a month and X Premium+ at $40. Speed comes at the cost of care: Grok leans personable over precise, and its citation reliability has not been independently proven, so verify anything that matters. Our Grok versus ChatGPT comparison has the specifics. ## Also Worth Knowing: Chinese Models, Vibe, and No-Login Picks A few more fill specific gaps rather than trying to replace ChatGPT wholesale. **Chinese hosted assistants** deserve a mention beyond DeepSeek. Kimi from Moonshot and Qwen Chat from Alibaba are both free, capable, and strong on reasoning and long context, and GLM from Zhipu is in the same tier. Treat them the way you treat DeepSeek: genuinely good models with the same data-processed-in-China caveat, so keep sensitive work off them. **Mistral's Vibe** (formerly Le Chat) is the European option, built by a French lab with a privacy-leaning stance and open-weight models underneath. It is the pick if data sovereignty or an EU-based provider matters to you, and it is multilingual by default. The free tier is usable, and Vibe Pro runs $14.99 a month. **For images**, none of the above is your first stop. Gemini generates and edits images natively, and ChatGPT's own image tools are still among the best, so image work is one job where staying put or switching to Gemini beats the rest of this list. **No-login answers** are their own job, separate from "free." DuckDuckGo's AI Chat and HuggingChat both give you a usable assistant without an account or a saved history, which is the fastest path to a quick answer you do not want logged. Meta AI covers the same casual use inside WhatsApp, Instagram, and Facebook. **Model aggregators** like Poe hand you many models under one subscription, so you can switch between Claude, GPT, and others in a single interface instead of paying for each. That is a practical hedge against betting on the wrong provider. ## Open-Source and Local: Run a ChatGPT Alternative on Your Own Machine If privacy is the reason you are leaving, the strongest answer is to run a model on your own computer, where nothing you type ever leaves the machine. This is the option most listicles skip, and it is more approachable than it sounds. Tools like [Ollama](https://ollama.com) and LM Studio install with a single step and let you download and run open models such as Llama, Qwen, and DeepSeek locally. The software is free. Your only cost is hardware and setup, and this is where the common myth breaks down: you do not need an expensive GPU to start. A small model runs on a normal laptop CPU, though expect weaker answers and slower output at that size, and you only need serious hardware when you want the largest models at full speed. Local models are also the real answer for anyone leaving because of over-refusal. Open-weight models are easy to run with far fewer guardrails than the hosted assistants, which is a fit for unfiltered research and a liability for anything brand-facing, so match the freedom to the use. Beyond that, a local model on your own laptop will not match the frontier hosted models on the hardest tasks, and you own the upkeep. But for private drafting, offline use, and anything you refuse to send to a vendor, it is the one option here that keeps your data entirely yours. ## How to Switch Without Regret The biggest mistake switchers make is treating this like a divorce, picking one tool and going all in. Route by task and keep a second option a click away. That backup also protects you from the thing that actually burns people. That thing is model deprecation. When OpenAI retired GPT-4o, users who had built prompts and workflows around it lost the exact behavior they relied on overnight. Any single provider can change or kill the model under you, so a setup you can move quickly beats loyalty to one brand. Tools that connect to your files and apps through open standards like the [Model Context Protocol](https://modelcontextprotocol.io) make that portability easier, because your setup is not welded to one vendor. Before you commit, check the data and training policy, because that is often the real switch trigger. By default, [OpenAI uses your ChatGPT conversations to improve its models](https://openai.com/policies/how-your-data-is-used-to-improve-model-performance/) unless you turn off "Improve the model for everyone" in Data Controls. Every alternative has its own answer to this question, and a privacy-leaning tool like a local model or Vibe is only worth switching to if its policy is actually better, so read it. One practical step that is easy to miss: you can bring your history with you. ChatGPT exports your data from Settings, and while nothing carries over automatically between providers, the workaround is quick. Claude and Gemini both accept file uploads, so you can feed that export straight in; aggregators like Poe and local tools let you paste the key parts as context. A few minutes of setup makes a switch feel far less like starting over. ## ChatGPT Is Now One of Several Answer Engines The way people look for a ChatGPT alternative has changed. Fewer search for a generic replacement; more search for Claude, Gemini, or Perplexity by name, and Google's AI Overviews answer many comparison questions before anyone clicks. The category did not consolidate around one winner. It spread into a handful of assistants people use in parallel. That spread matters beyond your own chat window. If you run a website or a brand, your customers are now asking questions across all of these tools, and each one decides on its own which sources to cite. Being mentioned by ChatGPT tells you almost nothing about whether Gemini, Perplexity, or Copilot surface you. Across the sites we scan for AI visibility at geotoolbox, the same page rarely shows up consistently across every engine, which leaves you under-cited in the ones you ignore. That scan is a sampled snapshot rather than a live crawl, but the pattern holds run after run. This is the shift behind the switching. The same forces pulling users toward alternatives are splitting where your audience gets its answers, so getting cited by ChatGPT is now one lane of a wider job. If that is your concern, our guides to AI visibility and optimizing for AI search cover how to earn citations across engines rather than one. ## Frequently Asked Questions ### Which AI is better than ChatGPT? It depends on the job. Claude is better for writing and long-form reasoning, Perplexity is better for researching with sources, and Gemini is better if you work inside Google. No one tool wins across the board, and crowd-ranked leaderboards like [LMArena](https://lmarena.ai/leaderboard) shift month to month, so pick by task rather than by an overall score. ### What is the best free ChatGPT alternative? Gemini, DeepSeek, and Meta AI all have strong free tiers, and Gemini is free inside Google Search. The idea that free tools are worse is out of date. For everyday questions, a free tier from any of these matches paid ChatGPT closely enough that most people never need to upgrade. ### Is there a ChatGPT alternative with no login? Yes. DuckDuckGo's AI Chat and HuggingChat both let you use an assistant without creating an account, and Meta AI works inside apps you already have signed into. These are the fastest options when you want an answer without a saved history. ### Which ChatGPT alternative is best for coding? Claude and DeepSeek are the two developers reach for most: Claude for reasoning through a large codebase, DeepSeek for strong results at a low cost. If you code inside an IDE, Copilot's in-editor suggestions are the tighter fit. ### Is there an open-source ChatGPT alternative? Yes. DeepSeek, Meta's Llama, and Alibaba's Qwen all release open-weight models you can run yourself using tools like Ollama or LM Studio. That gives you a private, offline assistant with no subscription, at the cost of some setup. ### Which AI does Elon Musk use? Grok, built by his own company xAI and wired into X. It is the alternative to reach for when you want real-time takes from the live feed, though its accuracy on careful tasks is less proven. ### Can I move my ChatGPT history to another app? Partly. You can export your ChatGPT data from Settings, then paste the important context or upload that export into the new tool to rebuild its working memory. There is no one-click transfer between providers yet, but a few minutes of setup carries most of what matters. ### Who is the biggest rival of ChatGPT? By reach, Google Gemini, since it ships free inside Search and Workspace to billions of users. By head-to-head capability, Claude is the one most reviewers put level with or ahead of ChatGPT on writing and reasoning. Perplexity is the biggest rival for search-style questions specifically. ### Is ChatGPT still the best AI? It is still the most popular all-rounder, but "best" is now task-specific. For cited research Perplexity wins, for long-document work Claude wins, and for anyone inside Google or Microsoft the built-in assistant often wins on convenience. ChatGPT is a safe default, and no longer the automatic first choice. ## Pick the Right Tool for the Job, Not the Loudest One The takeaway is simple: there is no single best ChatGPT alternative, only a best one for the job in front of you right now. Route by task, keep a backup so a retired model never strands you, and re-check prices before you subscribe, because they move faster than the articles quoting them, as the dead Copilot Pro tier proves. The wider version of this problem is not which assistant you use, but which ones your customers use, and whether they find you inside any of them. That is the job we built geotoolbox for. If you want to see how ChatGPT, Gemini, Perplexity, Claude, and the rest actually cite your site, our domain overview shows where you show up across every major answer engine and where you are missing. ## Sources - Microsoft launches 365 Premium for consumers, retires Copilot Pro - PCWorld, Mark Hachman, Oct 1 2025 - `pcworld.com/article/2925962` - How your data is used to improve model performance - OpenAI - `openai.com/policies/how-your-data-is-used-to-improve-model-performance` - Microsoft 365 Copilot plans and pricing - Microsoft - `microsoft.com/en-us/microsoft-365-copilot/pricing` - About Microsoft Copilot Pro (retirement + Aug 1 2026 sunset) - Microsoft Support - `support.microsoft.com/en-us/microsoft-365-copilot/about-microsoft-copilot-pro` - Ollama, run open models locally - Ollama - `ollama.com` - LMArena model leaderboard - `lmarena.ai/leaderboard` - Model Context Protocol - `modelcontextprotocol.io` --- ## How ChatGPT Search Works > The model doesn't browse the web. A retrieval tool fetches live pages, and it answers with citations. How ChatGPT Search works and which index it uses. - Canonical: https://geotoolbox.ai/blog/how-chatgpt-search-works - Published: 2026-07-23 · Updated: 2026-07-23 Ask ChatGPT about today's news and it answers with live links. Ask it the same question with search off and it either shrugs or makes something up. Same chatbot, completely different behavior. The difference is a feature called ChatGPT Search, and almost everything people get wrong about it comes from one misunderstanding of how it actually works. So here is the mechanism, start to finish: when it searches, which index it pulls from, when it decides to search at all, and why its answers differ from a Google results page. ## The Model Doesn't Browse. A Tool Does. The language model at the center of ChatGPT cannot open a web page. It generates text, one token at a time, from patterns it learned in training. That is the whole job. It has no browser, no network connection, no way to fetch a URL on its own. For the full picture of what the model itself does, see [how ChatGPT works](https://geotoolbox.ai/blog/how-does-chatgpt-work). So when ChatGPT "searches," something else does the searching. A separate retrieval tool runs the query, pulls back web pages, and drops their text into the model's context, the same place your prompt goes. The model then reads that text as if you had pasted it in yourself, and writes an answer that cites the sources. Put another way: large language models generate text, they do not browse, so the browsing is delegated to a tool and the model only ever sees the text that tool hands back. This pattern has a name: retrieval-augmented generation, or [RAG](https://geotoolbox.ai/blog/what-is-rag). It is the single fact that explains every quirk in the rest of this article. The model only knows what the tool handed it. If the tool retrieves the wrong page, or a thin snippet, or nothing at all, the answer degrades, and the model rarely tells you that happened. ## How a ChatGPT Search Runs, Step by Step A single search is a short pipeline. OpenAI has not published a full spec, so this is the widely-understood behavior, pieced together from its documentation and independent testing. Roughly: 1. **Decide whether to search.** An internal classifier looks at your prompt and judges whether it needs fresh information. A question about a breaking event or a current price triggers a search. A timeless question ("what is photosynthesis") gets answered straight from training. You can also force a search yourself with the search tool, the globe icon, or by starting a message with a slash command. 2. **Rewrite the query.** Your prompt is not sent to a search engine verbatim. The system rewrites it into one or more optimized queries, sometimes several at once, a technique called [query fan-out](https://geotoolbox.ai/blog/query-fan-out). What you typed and what it actually searched for are often different, which is why the results can surprise you. 3. **Retrieve from an index.** Those queries run against a web index and come back as a ranked list, the familiar page of blue links. At this stage the tool sees only metadata for each result: an internal ID, the title, the URL, a short snippet, and a last-updated date. No body text yet. 4. **Read the promising pages in chunks.** For results worth opening, the tool fetches the page and reads it in pieces rather than all at once. Dan Petrovic's [reverse-engineering of the process](https://dejan.ai/blog/how-gpt-sees-the-web/) describes a "sliding window" that jumps around the document, strips out design and scripts, and pulls short passages of plain text. It does not load the whole page, and the exact chunk sizes are not something OpenAI has published, so treat the specifics as observed rather than official. 5. **Synthesize and cite.** Finally the model combines those passages with its own trained knowledge and writes a conversational answer, attaching inline citations to the pages it leaned on. The takeaway from step 4 matters for anyone publishing on the web: if consumer search behaves like Petrovic's tests, the tool pulls fragments rather than whole pages, and the earliest windows are what capture the top of your content. It rarely holds a full article in view at once.
![A five-step flow diagram of how ChatGPT search works: the model decides whether to search or answer from training memory, rewrites the prompt into fan-out queries, retrieves ranked links from a Bing and OpenAI index using metadata only, reads short chunks of the top pages with a sliding window, then synthesizes one answer with inline citations.](/blog/how-chatgpt-search-works/chatgpt-search-retrieval-pipeline.png)
ChatGPT Search is retrieval-augmented generation: a separate tool searches and reads, then the model writes the cited answer.
## Which Search Index Does ChatGPT Use? Bing, OpenAI's Own, or Both Short answer: a hybrid, and OpenAI has not published the recipe. ChatGPT pulls from a Bing-class third-party index and from OpenAI's own web crawl. Both sources are confirmed; how they are weighted is not. What Google is not, on the evidence available, is a disclosed index partner, though it turns out to matter at the edges. Two pieces of infrastructure sit behind that. OpenAI's [own crawler documentation](https://developers.openai.com/api/docs/bots) lists three separate bots, and only one of the [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) feeds search:
CrawlerWhat it doesFeeds ChatGPT Search?
OAI-SearchBotCrawls the web to surface sites in ChatGPT's search featuresYes. Block it and your content won't be cited in Search answers, though the URL can still show as a plain link
GPTBotCollects content that may be used to train the modelsNo. It trains the models; it does not control Search eligibility
ChatGPT-UserMakes user-initiated page visits for some ChatGPT questions and actionsNo. Not the automatic Search crawler, and not used to decide what appears in Search
The Microsoft partnership means Bing's index has been a primary source from early on, and the fact that OAI-SearchBot exists at all is consistent with OpenAI also building its own cache or index rather than renting one outright. So where does Google come in? Not as a partner, but analysts keep catching it at the edges. Lily Ray [noted](https://www.linkedin.com/posts/lily-ray-44755615_it-seems-like-chatgpt-was-heavily-reliant-activity-7483170448348868608-CZBz) that ChatGPT looked "heavily reliant on Google's search results" until recently, and now appears to be "leaning on Bing results more heavily again, while building its own cache/index." Aleyda Solis [documented](https://www.searchenginejournal.com/chatgpt-appears-to-use-google-search-as-a-fallback/552089/) ChatGPT returning a snippet that matched Google's result for a page Bing had not indexed, which reads like a Google fallback when Bing comes up short. So the honest picture is fluid: Bing-class retrieval plus a growing OpenAI index as the primary sources, Google surfacing as an occasional fallback, and the exact mix undisclosed. The infrastructure data backs the "building its own index" part. A Botify and Nectiv [analysis of about 7 billion log files](https://www.botify.com/blog/openai-tripled-web-crawl) found OpenAI's crawl tripled after August 2025, with OAI-SearchBot activity rising 3.5x, enough that its search crawler now logs slightly more requests than its training crawler. ## When ChatGPT Searches vs Answers From Memory Not every question triggers a search, and this is the source of most complaints about accuracy. The classifier from step one makes the call: current or fast-moving topics get retrieval, timeless ones get answered from training data. That matters because training data has a cutoff. Ask about something that happened after the model's cutoff and, if the classifier decides not to search, you get a confident guess instead of a fact. The catch is that ChatGPT does not clearly flag which mode produced an answer. A wrong "fact" that a user blames on ChatGPT Search was often never a search at all, just the model answering from stale memory. The fix is to remove the guesswork. If freshness matters, force the search rather than hoping the classifier does. Turn on the search tool, or watch for the "Searching the web" status and the source links that only appear when retrieval actually ran. For heavier questions, Deep Research runs the same retrieval loop many times over, reading dozens of pages before it answers, which is slower but far more thorough than a single search. ## ChatGPT Search vs Google: What's Actually Different They are built for different jobs. Google hands you a ranked list of links and lets you do the reading and judging. ChatGPT Search does the reading for you and hands back one synthesized answer with citations. Everything else follows from that split. For how this plays out across engines, see our breakdown of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work).
DimensionChatGPT SearchGoogle Search
OutputOne written answer with a few inline citationsA ranked list of links you scan yourself
Sources readA handful of pages, chosen for youThousands ranked; you pick which to open
Follow-upsConversational, keeps contextEach search starts fresh
PersonalizationRough location and memory; less location-tunedPersonalized and location-aware
Best forSynthesis, comparisons, "explain this"Navigational, local, and transactional queries
Because it reads only a few pages, ChatGPT Search is strong when you want a synthesized take and weak when you want breadth or a specific site. It leans on rough, IP-level location rather than the tuned local results Google gives you, so "best coffee near me" is still a job Google does better. And the raw scale gap is real: in the same Botify dataset, Google's crawlers logged roughly 18.2 billion crawl events in the final month measured, against 887 million from OpenAI's two crawlers combined, more than 20 to 1. That is crawl activity, not searches, but it is a reminder that ChatGPT Search, growing fast as it is, remains a small slice of how the web gets crawled and read. Perplexity takes yet another approach to the same problem, which we cover in [ChatGPT vs Perplexity](https://geotoolbox.ai/blog/chatgpt-vs-perplexity). ## Is ChatGPT Search Free? Do You Need to Log In? Yes, it is free, and no, you do not need an account. Access opened up in stages. OpenAI launched ChatGPT Search on October 31, 2024 for Plus and Team subscribers. It reached logged-in free users in December 2024, and by February 2025 OpenAI dropped the login requirement entirely, so you can search from a signed-out session. Paid plans do not change what the feature is; they raise the ceiling, with higher usage limits and access to Deep Research. If you have seen the claim that ChatGPT Search is paid-only, it is out of date. The retrieval mechanism is identical whether you pay or not. ## Can You Trust It? Where ChatGPT Search Gets Things Wrong The citations look authoritative. They are not always accurate. A Columbia Journalism Review study that [tested eight AI search engines](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php) across 1,600 queries found they identified the source of a quote incorrectly more than 60% of the time. ChatGPT Search was wrong in 134 of its 200 tests (67%), giving a partly or fully incorrect source and rarely flagging any uncertainty. The mechanism explains why. The tool screens results on metadata alone, then reads winners in disconnected chunks, then the model writes fluent prose around whatever it pulled. A citation is generated to look right, not guaranteed to be right. There is also a reachability trap: when OAI-SearchBot cannot reach the original page, ChatGPT may cite a scraper or aggregator that copied the content instead of the source itself. We go deeper on this in [how ChatGPT cites sources](https://geotoolbox.ai/blog/chatgpt-citations). The practical rule: a plausible-looking link is not verification. For anything that matters, click through and confirm the source actually says what the answer claims. ## What This Means If You Want ChatGPT to Find You Before any content tactic, one prerequisite decides everything: OAI-SearchBot has to be able to reach and read your pages. If it is blocked in robots.txt, or your content only renders after JavaScript the crawler will not run, you are not eligible to be cited, no matter how good the writing is. And because the tool reads in chunks from the top, a clear answer in your first few paragraphs beats the same answer buried under an introduction. That reachability layer is the most common gap we see when we scan sites for AI visibility: the owner blocked or broke a crawler without meaning to, and it is invisible until you check the logs. Note that this is a different job from getting cited once you are reachable, which involves reputation, structure, and freshness. We keep the two separate on purpose, and the get-cited playbook lives in our guide to [SEO for ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt). If you are not sure whether ChatGPT's crawler can even see your site, that is the first thing worth checking. Our [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) tests whether OAI-SearchBot and the other AI bots are allowed to reach your pages, so you can fix the plumbing before you worry about the content. ## Frequently Asked Questions ### Can I use ChatGPT for search? Yes. ChatGPT often searches automatically when a question needs current information, and you can force it any time with the search tool or globe icon. It returns a written answer with links rather than a list of results. ### Is ChatGPT Search free? Yes. It is free and, since February 2025, works without an account. Paid plans raise your usage limits and add Deep Research, but the core search feature is the same on the free tier. ### Does ChatGPT use Bing or Google? Mostly Bing, plus OpenAI's own crawler, OAI-SearchBot. Google is not a disclosed index partner, though analysts have caught ChatGPT falling back to Google results when Bing lacks a page. OpenAI has not said how the sources are weighted. ### Is ChatGPT Search better than Google? They do different jobs. ChatGPT is better for synthesis, explanations, and follow-up questions. Google is better for navigational, local, and transactional searches where you want to pick from many sources yourself. ### How do I turn on search in ChatGPT? Click the globe or search icon before sending your message, or start typing with the search tool selected. In many cases ChatGPT triggers a search on its own and shows a "Searching the web" status while it does. ### Can other people see my ChatGPT searches? Not by default. A search sits in your own chat history, or your session if you are signed out, and is not posted anywhere public. A couple of caveats: you can choose to share a chat, and to actually run the search your query (rewritten) and a rough location are sent to the search provider. ## Sources - Introducing ChatGPT search - OpenAI, October 2024 - `openai.com/index/introducing-chatgpt-search/` - Overview of OpenAI Crawlers (OAI-SearchBot, GPTBot, ChatGPT-User) - OpenAI - `developers.openai.com/api/docs/bots` - OpenAI Has Tripled Their Crawl of the Web: 7B+ Log Files Analyzed - Botify and Nectiv, April 2026 - `botify.com/blog/openai-tripled-web-crawl` - ChatGPT is leaning on Bing again while building its own index - Lily Ray, LinkedIn, 2026 - `linkedin.com/posts/lily-ray-44755615_it-seems-like-chatgpt-was-heavily-reliant-activity-7483170448348868608-CZBz` - ChatGPT Appears to Use Google Search as a Fallback - Aleyda Solis, Search Engine Journal, July 2025 - `searchenginejournal.com/chatgpt-appears-to-use-google-search-as-a-fallback/552089/` - How GPT Sees the Web (sliding-window retrieval research) - Dan Petrovic, DEJAN - `dejan.ai/blog/how-gpt-sees-the-web/` - AI Search Has a Citation Problem - Columbia Journalism Review, Tow Center, March 2025 - `cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php` - ChatGPT Search - OpenAI Help Center - `help.openai.com/articles/9237897-chatgpt-search` --- ## How to Turn Off AI Overviews: What Actually Works in 2026 > No official setting turns off Google AI Overviews. What works in 2026: the Web filter, udm=14, extensions, mobile fixes, and the real opt-out for site owners. - Canonical: https://geotoolbox.ai/blog/how-to-turn-off-ai-overviews - Published: 2026-07-23 · Updated: 2026-08-05 If you're searching for how to turn off AI Overviews, here's the part most guides bury: there is no off switch, and there never has been. What exists instead is a set of workarounds with very different lifespans, from Google's own Web filter to extensions that only paper over the box. Some of those workarounds hold up in 2026, several broke since the advice you read last year, and one genuine opt-out now exists if you run a website. The difference between them comes down to what each one actually does under the hood. If you first want the mechanics of the feature itself, we cover [what AI Overviews are](https://geotoolbox.ai/blog/what-are-google-ai-overviews) and how Google builds them separately. ## Can You Turn Off AI Overviews Completely? No. There is no setting in your Google account, in Search preferences, or anywhere else that disables [AI Overviews](https://geotoolbox.ai/glossary/ai-overviews) globally. Google's own help documentation is unambiguous: AI Overviews are [a core Google Search feature](https://support.google.com/websearch/answer/14901683), "like knowledge panels," and "Features cannot be turned off." That sentence ends most of the advice you'll find in older articles. The Search Labs toggle that people still recommend never disabled AI Overviews outside of Labs experiments, and Google has since retired it for most accounts and regions as the feature left its experimental phase. What remains in Labs today is an "AI in Search" opt-out that only stops **experimental** AI features from reaching you early. The AI Overviews you see in normal results stay exactly where they are. Incognito mode, logging out, and digging through account settings don't help either; frustrated users have tried all three, and the overview stays. (Gemini inside Gmail and Docs has its own Workspace settings; different feature, different article.) Google hasn't said why there's no off switch. But AI Overviews are its answer to ChatGPT and Perplexity, and they now run in [more than 200 countries and territories](https://blog.google/products-and-platforms/products/search/ai-overview-expansion-may-2025-update/) and more than 40 languages. A feature this strategic is unlikely to ever get a checkbox. The demand for an exit isn't abstract, either. The feature's early runs produced answers like the infamous glue-on-pizza suggestion, [a result Google itself acknowledged](https://blog.google/products/search/ai-overviews-update-may-2024/) while defending overall accuracy, and the skepticism stuck. You cannot turn AI Overviews off. You can stop seeing them, which is a different thing, and for most people just as good. Every method below does one of three jobs: it requests a results page that never includes the overview, it hides the overview after it loads, or, if you own a website, it pulls your content out of the overviews Google shows other people.
![Three cards comparing the ways to turn off Google AI Overviews: extensions hide the box after it is generated, the Web filter and udm=14 request a results page without it, and the Search Console toggle lets site owners exclude their content at the source.](/blog/how-to-turn-off-ai-overviews/three-ways-to-turn-off-ai-overviews.png)
Hiding, filtering, and opting out are three different jobs, and the clean opt-out for site owners started in the UK and is spreading, still partially.
## The Web Filter and udm=14: The Closest Thing to an Off Switch Google ships an AI-free version of its own results. It's called the **Web filter**, it sits under the "More" tab after you run a search, and it returns plain blue links with no AI Overview, no shopping modules, and no panels. Google added it in May 2024, the same month AI Overviews launched in the US. Clicking "More > Web" on every single search is tedious. That's where **udm=14** comes in. The Web filter is just a URL parameter: append `&udm=14` to any Google search URL and you land on the text-only page directly. Tedium's Ernie Smith documented the trick days after the filter shipped in [a post that went viral](https://tedium.co/2024/05/17/google-web-search-make-default/), then put up [udm14.com](https://udm14.com) so anyone could use it without touching a URL. A search on udm14.com is a Google search with the parameter pre-applied; it does route through his third-party page first, so if that bothers you, the custom-engine setup below keeps everything between you and Google. Here's what matters about udm=14 compared with every other trick in this article: the page arrives without an AI Overview in it. You are not hiding the feature on your side of the screen, you are requesting the version of Google that doesn't include it. The trade-off is that the Web filter strips more than AI. You lose shopping results, knowledge panels, video carousels, and most local map packs. For research queries that's usually a feature. For "restaurants near me" it isn't, so keep a normal Google tab in your routine for local searches. One setting to leave alone: the **Web Guide** experiment in Labs re-adds AI summaries to this same Web tab (more on that below). ## Make AI-Free Google Your Default Search Engine The durable desktop fix is a custom search engine that applies udm=14 to everything you type in the address bar. Set it once and AI Overviews disappear from every search you start in the address bar. In Chrome: 1. Open `chrome://settings/searchEngines` (or Settings > Search engine > Manage search engines and site search). 2. Under **Site search**, click **"Add"**. 3. Fill in the three fields: Name: `Google (Web)` - Shortcut: `@web` - URL: `{google:baseURL}search?udm=14&q=%s` 4. Click **"Add"**, then open the three-dot menu next to your new entry and select **"Make default"**. From now on, every address-bar search goes through the Web filter. Type the query, hit Enter, get links. Edge has the same flow under Settings > Privacy, search, and services > Address bar and search. Brave mirrors Chrome's settings almost exactly. In Firefox, add a custom search engine with `https://www.google.com/search?udm=14&q=%s` under Settings > Search. Safari on Mac allows no custom search engines at all; use a udm14.com bookmark or the query tricks below. Two caveats. First, this filters your results; the AI Overviews feature itself stays enabled on Google's side, and searches started on google.com directly (rather than the address bar) still show overviews. Second, Google has never documented udm=14 as a stable public interface. It has worked continuously since May 2024, but it works because Google lets it. ## Extensions Hide AI Overviews (but Don't Stop Them) Browser extensions are the most popular route, and the most misunderstood one. An extension does not prevent the AI Overview from being created. Whenever a query triggers one, Google still generates the answer server-side; the extension applies CSS that hides the box before you see it. You get a cleaner page; the answer underneath was still generated. That distinction matters more than it sounds. If your objection to AI Overviews is the interface clutter, hiding works fine. If your objection is the feature existing at all, an extension changes nothing underneath. The best-known option is **Bye Bye, Google AI**, built by Tom's Hardware's Avram Piltch, who published it after arguing that AI Overviews take publisher content [without sending traffic back](https://www.tomshardware.com/how-to/block-google-ai-overviews). It's open source, works in Chrome and Firefox, and can also hide the AI Mode tab, "People also ask," and other modules you may not want. A separate extension, Hide Google AI Overviews, does the narrower version of the same job. And if you already run uBlock Origin, custom cosmetic filters can hide the overview too; treat that as the advanced route, because filters target Google's current page markup and break silently when it changes. The catch is durability. Extensions can break whenever Google changes the page markup they target, and they get repaired on the developer's schedule, not yours. Chrome also [finished disabling Manifest V2 extensions in 2025](https://developer.chrome.com/docs/extensions/develop/migrate/mv2-deprecation-timeline), which permanently killed older blockers that were never rebuilt for Manifest V3. If a hidden overview suddenly reappears, check the extension before assuming Google patched anything. One safety note: install these only from the official Chrome Web Store or Firefox Add-ons listings, and read the permissions. An extension that can rewrite your Google results page can read everything on it. Skip anything that asks for more than that job requires. ## Query Tricks: The -AI Operator and Friends For one-off searches, there's a faster route: append a negative operator to your query. Searching `best running shoes -AI` usually returns a page with no AI Overview at all. Here's why. A minus operator tells Google to exclude results containing that term, and queries carrying exclusion operators generally don't trigger the overview. Google doesn't document this behavior anywhere, and the exclusion is real: `best running shoes -AI` also drops pages that mention AI, so on AI-adjacent topics the operator distorts what you get back. The effect is per-query, so it suits the occasional search on a machine you don't control. For daily use, set a default engine instead. Two related tricks deserve a status update, because both circulated widely and neither aged well. Adding profanity to queries suppressed AI Overviews for a while, but it distorts what Google returns, can pull results toward content you didn't want, and has been patchy since mid-2025; skip it. Voice search sometimes skips the overview and sometimes doesn't, so it was never reliable enough to recommend. That's the pattern with query tricks generally: whack-a-mole. Google adjusts, a trick dies, a new one circulates on Reddit. The methods that persist are the ones built on Google's own Web filter; Google can quietly patch a loophole, but removing a documented feature is a much bigger step. ## How to Turn Off AI Overviews on iPhone and Android Mobile is where the frustration concentrates. The ceiling first: **the Google app ships no AI-free setting**, and neither does the search widget on your home screen. Every workaround below changes what happens in a browser; inside the app and widget, the only lever is typing the `-AI` operator into the query itself. With that said, here is what works on a phone. **Firefox (Android and iOS)** is the one mobile browser that lets you add a custom search engine by hand. Open Settings > Search > Default Search Engine > Add search engine, name it `AI-free Google`, and paste `https://www.google.com/search?udm=14&q=%s` as the URL. Set it as default and your mobile address bar behaves like the desktop fix above. **Chrome on Android** won't let you type in a custom engine, but it will let you pick one it has "seen." Visit [tenbluelinks.org](https://tenbluelinks.org) once, then open Settings > Search engine in Chrome: a "Google Web" option appears under recently visited engines. Select it and address-bar searches route through the Web filter. **Safari on iPhone** allows neither trick; Apple restricts search engines to a fixed list. Your options there are a bookmark and the `-AI` operator. Save `https://udm14.com` (or a udm=14 search URL) to your home screen and run searches from it, or type the minus operator when it matters. If that all sounds like more friction than switching apps, switch apps. DuckDuckGo and Brave both ship mobile search apps where AI answers are optional settings rather than fixtures, and for plenty of people that is the practical mobile answer. ## Which Method Should You Use? Every approach trades something. This is the whole picture in one place:
MethodWorks onPersistenceWhat you loseEffort
Web filter (More > Web)Desktop + mobile browsersOne search at a timeShopping, panels, local packsNone
-AI operatorEverywhere, including the Google app (undocumented, not guaranteed)One search at a timeResults mentioning the excluded term drop outNone
udm=14 bookmark (udm14.com / tenbluelinks.org)Any browserAs long as you start thereSame as Web filterOne-time
Custom default search engineDesktop browsers + Firefox mobilePermanent for address-bar searches, until Google changes udm=14Same as Web filter; local searches need a detourOne-time setup
Extension (Bye Bye, Google AI etc.)Desktop Chrome/Firefox/EdgeUntil Google changes the pageNothing visible; overview still generatedOne-time + upkeep
Different search engine (DuckDuckGo, Brave, Startpage, Ecosia, Kagi)EverywherePermanentGoogle's index and result quality on some nichesOne-time
The short version: pick the one permanent option your platform allows, and fall back to per-query tricks everywhere else. The per-device setup is in the closing section. ## Can You Turn Off AI Mode Too? No, and the features are worth separating, because 2026 blurred them. AI Overviews are the summary box on top of normal results. **AI Mode** is a full conversational search tab that Google has been pushing hard; it [passed a billion monthly users](https://blog.google/innovation-and-ai/sundar-pichai-io-2026/) by mid-2026. You can ignore the AI Mode tab, but there is no setting that removes it, and Google has been testing flows that route some queries into it directly. Everything in this article about udm=14 and the Web filter applies: a Web-filtered results page carries no AI Mode answer. If you want the deeper picture of what AI Mode changes (and what it means if you run a website), we covered it in our guide to [Google AI Mode SEO](https://geotoolbox.ai/blog/google-ai-mode-seo). One warning for Web tab users: Google's **Web Guide** experiment in Search Labs reorganizes the Web tab with Gemini and reintroduces AI summaries below the links. If you enabled Web Guide out of curiosity and your "clean" tab grew AI content back, that toggle is why. Turn it off in Labs and the plain Web tab returns. ## Site Owners: Opting Your Content Out of AI Overviews Everything above is about not seeing AI Overviews. If you run a website, you have the opposite problem: your content appearing inside them, answering the searcher's question before they ever reach your page. The concern is measurable. [Pew Research Center's analysis](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) of 900 US adults' browsing found that users clicked a traditional result on just 8% of visits when an AI summary was present, versus 15% without one, and clicked a source cited inside the summary on only 1% of visits. This is not a niche encounter: 58% of the panel hit at least one AI summary in March 2025 alone. For years the only controls were blunt. Now there are five levers people reach for, and one of them does not do what most people think:
ControlWhat it actually doesSide effects
Search Console "Search generative AI" toggleExcludes your property from AI Overviews, AI Mode, and generative AI in Discover; usually takes effect in 1-2 days, longer for cached contentStarted UK-first, now partially global — a subset of publishers under a CMA order dated June 3, 2026, spreading to some non-UK accounts since July 2026 with no completion date; per verified property, with inheritance for child properties
nosnippet robots tagBlocks your text from AI OverviewsAlso kills your normal search snippets - the description under your blue link
data-nosnippet attributeExcludes marked page sections onlySurgical, but you must wrap the right HTML
max-snippet:[n]Caps how much text Google may quote; low values reduce AI use but don't guarantee exclusion (max-snippet:0 equals nosnippet)Shortens regular snippets too
Blocking Google-Extended in robots.txtStops Gemini training and grounding use of your contentDoes not touch AI Overviews. A common misconfiguration in the wild
The new [Search Console control](https://support.google.com/webmasters/answer/16908024) is the first real opt-out: Google states it is not a ranking signal and takes effect within a day or two. But it exists because the UK's Competition and Markets Authority forced it, it [started with a subset of UK properties](https://ppc.land/google-gives-site-owners-a-toggle-to-exit-ai-overviews-and-ai-mode/) and has since spread to some non-UK accounts too, still a partial rollout with no completion date announced. Everyone else is still choosing between snippet directives that punish normal search visibility as the price of AI exclusion. Which is why the actual first step comes before any toggle or tag: find out whether AI Overviews cite you at all, and for which queries. The Pew numbers cut both ways: citations rarely turn into clicks, so staying in mostly buys you brand presence inside the answer, and opting out mostly costs you that presence. Either way, the decision deserves data. We built geotoolbox around exactly that measurement problem, and our guide to [getting cited in AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) covers the other direction, with [AI Overview tracking options](https://geotoolbox.ai/blog/ai-overview-tracker) compared separately. ## The Realistic Setup You cannot turn off AI Overviews. You can build a search routine where you never see them: a udm=14 default engine on desktop, Firefox with a custom engine or a udm14 bookmark on your phone, and the `-AI` operator for everything in between. Those hold up because they lean on Google's own Web filter instead of fighting the page. And if Google ever retires udm=14, the fallback is already in the comparison table: a search engine where AI answers are optional or absent. If you're on the site-owner side of this, the trade-off is starker: the one clean opt-out started UK-first and is now partially rolling out elsewhere too, and the other controls trade away snippet visibility in proportion to what you mark, from single sections with data-nosnippet to every snippet with nosnippet. Before you pull anything out of AI Overviews, look at where Google's AI actually mentions and cites your domain today. A [domain overview scan](https://geotoolbox.ai/features/domain-overview) in geotoolbox shows you that picture across engines, so the decision rests on numbers. ## Frequently Asked Questions ### Can I permanently turn off AI Overviews? No. Google's help documentation classifies AI Overviews as a core Search feature with no user setting to disable it. The closest thing to permanent is a custom default search engine with the udm=14 parameter, which routes every address-bar search to Google's AI-free Web filter. ### Does udm=14 still work in 2026? Yes. It has worked continuously since May 2024 because it's the URL form of Google's own Web filter, not an exploit. Bear in mind it strips shopping results, knowledge panels, and most local packs along with the AI, and Google has never promised the parameter is permanent. ### How do I turn off AI Overviews on iPhone? Safari won't accept custom search engines, and the Google app has no AI-free mode. Use Firefox for iOS with a custom search engine URL ending in `udm=14&q=%s`, keep a udm14.com bookmark on your home screen, or add `-AI` to individual queries. ### Did Google remove the AI Overviews opt-out? The Search Labs toggle people remember was an experiment control, not an opt-out; it never disabled AI Overviews in regular results, and Google has retired it for most accounts. The only formal opt-out that exists now is the Search Console exclusion control for site owners, which started UK-first and is now partially rolling out to accounts elsewhere too. ### Does blocking Google-Extended keep my site out of AI Overviews? No. Google-Extended controls whether Gemini can train on and ground against your content. AI Overviews eligibility follows normal Search indexing and snippet rules, so the controls that matter are nosnippet, data-nosnippet, max-snippet, or the Search Console toggle where available. ## Sources - AI Overviews in Google Search - Google Search Help - `support.google.com/websearch/answer/14901683` - Search generative AI control - Google Search Console Help - `support.google.com/webmasters/answer/16908024` - Google users are less likely to click on links when an AI summary appears in the results - Pew Research Center, July 22, 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/` - Does One Line Fix Google? - Tedium (Ernie Smith), May 17, 2024 - `tedium.co/2024/05/17/google-web-search-make-default/` - Google is killing the web with AI Overviews - I made an extension to block them - Tom's Hardware (Avram Piltch) - `tomshardware.com/how-to/block-google-ai-overviews` - Google gives site owners a toggle to exit AI Overviews and AI Mode - PPC Land, June 2026 - `ppc.land/google-gives-site-owners-a-toggle-to-exit-ai-overviews-and-ai-mode/` - Manifest V2 deprecation timeline - Chrome for Developers - `developer.chrome.com/docs/extensions/develop/migrate/mv2-deprecation-timeline` - AI Overviews: About last week - Google (Liz Reid), May 2024 - `blog.google/products/search/ai-overviews-update-may-2024/` - Expanding AI Overviews to more places and languages - Google, May 2025 - `blog.google/products-and-platforms/products/search/ai-overview-expansion-may-2025-update/` - Sundar Pichai's I/O 2026 keynote - Google, June 2026 - `blog.google/innovation-and-ai/sundar-pichai-io-2026/` - udm14.com - Ernie Smith's AI-free Google search shortcut - `udm14.com` - tenbluelinks.org - Web filter setup helper for mobile browsers - `tenbluelinks.org` --- ## What Are Google AI Overviews? How They Work > What Google AI Overviews are, how Gemini builds them with query fan-out, when they appear, how sources get cited, and what they mean for your site's traffic. - Canonical: https://geotoolbox.ai/blog/what-are-google-ai-overviews - Published: 2026-07-23 · Updated: 2026-08-05 Google AI Overviews are the AI-written answers that now sit above the blue links on a large share of Google searches. If you have searched for anything question-shaped in the past two years, you have probably seen one, whether you wanted it or not. ## What Are Google AI Overviews? **AI Overviews (AIOs) are AI-generated summaries that appear at the top of Google search results, synthesizing an answer from multiple web sources and linking to the pages the answer draws on.** They answer the query directly on the results page, usually at the very top and occasionally beneath sponsored results, and they are the most visible of Google's AI search features. Google [describes them](https://search.google/ways-to-search/ai-overviews/) as "a snapshot of key information about a topic or question with links so you can easily explore more on the web." The framing matters: Google positions the links as the point, publishers tend to see them as the consolation prize. A typical AI Overview has three parts: 1. **The generated summary.** A few sentences to several paragraphs of AI-written text answering the query, sometimes with bullets, images, or expandable sections. 2. **Citations.** Link cards and inline references pointing to the web pages Google offers as support for the answer. These are the "ranking positions" of the AI Overview era. 3. **A handoff to AI Mode.** A "Show more" or follow-up prompt that carries the query into [AI Mode](https://geotoolbox.ai/blog/what-is-google-ai-mode), Google's conversational search tab, with the context retained. An AI Overview is not a chatbot. You cannot argue with it, and it does not remember you between searches. It is a one-shot summary bolted onto the classic results page, which is exactly why it changed SEO instead of replacing it: the ranking systems underneath still decide which pages are even eligible to be cited. That question, who gets pulled into the answer and who gets left below it, is what [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) is about, and it is the lens for most of this article. ## From SGE to Today: How AI Overviews Evolved AI Overviews went from a hidden lab experiment to the default front door of Google search in two product cycles. The [Wikipedia entry on AI Overviews](https://en.wikipedia.org/wiki/AI_Overviews) keeps a running history; the short version:
WhenWhat happened
May 2023Google announces the Search Generative Experience (SGE) at I/O, an opt-in Search Labs experiment that put generated answers on top of results
May 14, 2024SGE graduates and launches as AI Overviews in the US, no opt-in required
Late May 2024Viral failures force Google to restrict the feature within weeks of launch
October 2024Expansion past 100 countries; ads begin appearing inside AI Overviews in the US
March 2025Gemini 2.0 rolls into AI Overviews; frequency more than doubles across 2025 as trigger thresholds loosen
May 2025Google I/O: AI Overviews reach 200+ countries and territories in 40+ languages; AI Mode launches in the US
2026AI Overviews and AI Mode fold closer together ("Show more" hands off with context); Search Console gets a generative-AI report in June; AI Mode passes one billion monthly users
One telling detail about how fast this moves: Google's own [product page for AI Overviews](https://search.google/ways-to-search/ai-overviews/) still lists availability at "over 120 countries and territories, and 11 languages," while its I/O 2025 announcement said 200+ countries and 40+ languages. Even Google's marketing copy cannot keep up with Google. You will also still find explainers claiming AI Overviews run on PaLM 2. That was the SGE-era model. Every current Google description points to Gemini, which is what the next section covers. ## How Do AI Overviews Actually Work? An AI Overview is built in a pipeline: expand the query, retrieve pages, generate a grounded summary, attach citations. Understanding the stages tells you where a site can win or lose a spot in the answer.
![Four-stage pipeline showing how Google builds an AI Overview, from query fan-out to citations.](/blog/what-are-google-ai-overviews/how-ai-overviews-work.png)
The four stages between a search and the generated answer: fan-out, retrieval, grounded generation, citation.
**Stage 1: query fan-out.** For complex searches, Google can decompose the query into multiple related sub-queries and run them in parallel. A search like "are AI overviews accurate" quietly becomes questions about error rates, incidents, and Google's fixes. This is [query fan-out](https://geotoolbox.ai/blog/query-fan-out), and it is why a page that answers a whole neighborhood of sub-questions gives retrieval more to grab than a page that answers exactly one. **Stage 2: retrieval.** The fan-out queries pull candidate pages from Google's regular search index. There is no separate "AI index." If a page is not indexed and eligible for a snippet, it cannot appear in an AI Overview, full stop. **Stage 3: grounded generation.** A customized version of Gemini writes the summary. Crucially, it is not answering from memory the way a raw chatbot does. Google says AI Overviews are ["built to only show information that is backed up by top web results"](https://blog.google/products-and-platforms/products/search/ai-overviews-update-may-2024/) and are integrated with its core web ranking systems. The industry name for this pattern is retrieval-augmented generation, and it is the same mechanic behind [most AI search engines](https://geotoolbox.ai/blog/how-does-ai-search-work). **Stage 4: citation.** The generated sentences get linked to the supporting pages, shown as inline references and link cards. Which passages get cited is a selection distinct from ranking, and it gets its own section below. The grounding is why AI Overviews are more reliable than a bare LLM answer, and the generation is why they still fail in ways a featured snippet never could. A snippet quotes a page verbatim; an AI Overview writes new text, so a bad retrieval or a satirical source becomes a confident wrong answer. ## When and Where Do AI Overviews Appear? AIOs trigger selectively, and the pattern is consistent across the major studies: question-form, informational, longer, and more complex search queries get them; navigational, local, news, and shopping queries mostly do not. Google's stated rule is that they appear when its systems decide a generated answer would be more useful than a list of links. How often is "often"? Here is where the numbers get messy, because every published figure measures something different:
StudyFigureWhat it actually measured
Pew Research Center (Mar 2025 data)18% of searchesReal behavior: 68,879 searches by 900 US adults; share of visited results pages that showed an AI summary
Semrush (Nov 2025)15.69% of keywords10M+ tracked keywords; the share swung from 6.49% in January 2025 to a 24.61% peak in July before settling
Ahrefs (May 2025)9.46% of keywords, 12.8%+ of volume590M keywords in its index; AIOs concentrate on high-volume queries, so the volume share runs higher than the keyword share
The figures disagree because the studies disagree on unit, sample, and date at once. A random keyword database, a commercially-tracked keyword set, and a panel of real humans sample different slices of search, and the feature itself more than doubled in frequency during 2025 while they measured, which is why even the two keyword-share studies land far apart. The still-bigger headline figures you may have seen, half of all searches and up, generally weight by search volume or come from single-vendor tracked query sets. Read any single "X% of searches have AI Overviews" stat with the methodology attached, or not at all. Our [state of AI search report](https://geotoolbox.ai/blog/state-of-ai-search-2026) tracks these numbers, and their contradictions, across engines. What the studies agree on is direction and skew. The [Semrush AI Overviews study](https://www.semrush.com/blog/semrush-ai-overviews-study/) found the informational share of AIO-triggering keywords falling as commercial queries jumped from 8.15% to 18.57% and navigational from under 1% to 10.33% in a year: the feature is expanding down the funnel. And [Ahrefs' analysis of 55.8 million AI Overviews](https://ahrefs.com/blog/insights-from-56-million-ai-overviews/), a separate study of the overviews themselves rather than its keyword index, confirmed they lean hard toward informational, longer, higher-volume queries while avoiding branded and local ones. If your traffic lives on question keywords, assume an AI Overview is or will be sitting on top of you. If you sell things or serve a local area, the exposure is real but smaller. ## How Do AI Overviews Choose and Cite Sources? Citation is not ranking. The cited sources inside an AI Overview overlap with the top organic results, but the overlap is partial and has been shrinking, which is why "we rank #1 and still aren't in the answer" is now a normal complaint. Two findings frame it. Grounding in top web results is Google's stated rule, but in practice retrieval reaches well beyond position one. [Ahrefs' 55.8M-overview dataset](https://ahrefs.com/blog/insights-from-56-million-ai-overviews/) shows citation coverage concentrating hard at the top of the web: the 50 most-cited domains take 28.90% of all AIO mentions, led by Reddit, Wikipedia, Quora, YouTube, and NIH.gov. Everyone else splits the rest, passage by passage. Reddit leading that list has context. Google has been surfacing more forum threads and firsthand perspectives in results since 2023, and it [signed a data deal with Reddit](https://www.tomsguide.com/ai/google-strikes-dollar60m-deal-with-reddit-for-ai-training-data-what-you-need-to-know) in early 2024, reported at about $60 million a year; Google has never said the deal touches citation selection. What the data does show is that the forum-content pullback after the 2024 incidents did not dethrone Reddit as the most-cited domain by mentions. The selection behaves like a contest between passages, not pages. Ranking asks "which page best deserves this position." Citation asks "which passage best supports this sentence of the answer." A page can win the first contest and lose the second by burying its answer, hedging it, or spreading it across sections. The practical bet: self-contained passages that state a fact, a number, and its context in one liftable block are the easiest material for the generation stage to use. That is a mechanical, fixable problem, and learning to optimize for it is a different discipline from classic on-page SEO. The tactics, from answer-first formatting to reachability, are covered in our guides to [getting cited in Google AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) and [getting cited by AI engines generally](https://geotoolbox.ai/blog/how-to-get-cited-by-ai). The one-line summary: you cannot buy your way in, you can only be the clearest source for one specific sentence of the answer. ## Are AI Overviews Accurate? Mostly, with failures that are rare in percentage terms and memorable in every other way. The feature's reputation was set in its first two weeks. In late May 2024, screenshots of AI Overviews [recommending glue on pizza](https://www.forbes.com/sites/roberthart/2024/05/31/google-restricts-ai-search-tool-after-nonsensical-answers-told-people-to-eat-rocks-and-put-glue-on-pizza/) and endorsing a rock a day went viral, sourced from a joke Reddit comment and an Onion article the model took literally. Google restricted the feature within days: less weight on forum content, guardrails on health queries, better detection of nonsensical prompts. Worth knowing: fact-checkers found a number of the viral screenshots were fabricated, which says something about the discourse around this feature too. Google's defense, in [its official post-mortem](https://blog.google/products-and-platforms/products/search/ai-overviews-update-may-2024/), is that many failures came from "data voids" (queries with little serious content to ground on) and that content-policy violations appeared in fewer than 1 in every 7 million unique queries, with accuracy it describes as on par with featured snippets. Independent spot-checks are less flattering than the 1-in-7-million framing suggests, because "not a policy violation" is a lower bar than "correct." When [WordStream tested AI Overviews on paid-search topics](https://www.wordstream.com/blog/google-ai-overviews) it knows cold, it found errors in roughly 26% of the answers it reviewed. An informal test in a single vertical, so it cannot estimate a platform-wide error rate, but it is consistent with what practitioners see: the summaries are weakest exactly where the underlying web content is thin, contested, or fast-moving. A subtler failure gets less attention: the answer can be broadly right while the page it cites does not support that specific sentence, so a visible source list is not the same thing as a verified one. The practical calibration: treat an AI Overview like a well-read stranger's summary. Fine for orientation, not for decisions with stakes, and always one click away from the actual source. That failure mode, a fluent answer grounded in the wrong or misread page, is the search-specific flavor of [AI hallucination](https://geotoolbox.ai/blog/ai-hallucinations). ## What AI Overviews Mean for Websites and Publishers The traffic impact is the hardest part to dispute: when an AI Overview answers the query, fewer people click anything. The cleanest behavioral evidence is [Pew Research Center's panel study](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) of 68,879 real searches: users clicked a traditional result on 8% of searches with an AI summary versus 15% without one, clicked a source link inside the summary on just 1% of them, and ended their session outright more often (26% vs 16%). Once a summary is present, clicking anything at all becomes the exception. Rank-tracking data says the same thing louder over time. Ahrefs' updated comparison of 300,000 keywords found the top-ranking page's average clickthrough rate is [58% lower](https://ahrefs.com/blog/ai-overviews-reduce-clicks-update/) when an AI Overview is present, based on December 2025 data, up from 34.5% in its April 2025 measurement. The modeled gap widened sharply between the two measurements, eight months apart, and the design is observational, so read it as the size of the correlation, not a proven causal split. The publisher fight has moved to court. In September 2025, Penske Media, the owner of Rolling Stone, Billboard, and Variety, [sued Google over AI Overviews](https://techcrunch.com/2025/09/14/rolling-stone-owner-penske-media-sues-google-over-ai-summaries/), the first major US news publisher to do so, arguing the summaries repackage its journalism and gut the clicks that fund it. Google's counter is that AI Overviews make search more helpful and "send traffic to a greater diversity of sites." Both things can be partially true, and neither pays a newsroom. Penske is not fighting alone. Edtech company [Chegg filed its own suit](https://www.cnbc.com/2025/02/24/chegg-sues-google-for-hurting-traffic-as-it-considers-alternatives.html) in early 2025 blaming AI Overviews for its traffic collapse, and a coalition of independent publishers [lodged an EU antitrust complaint](https://searchengineland.com/google-faces-eu-antitrust-complaint-over-ai-overviews-458123) in mid-2025, arguing sites cannot opt out of being summarized without leaving Google Search entirely. The honest framing for site owners: this is the [zero-click search pattern](https://geotoolbox.ai/blog/zero-click-searches) with a citation layer attached (and, on some commercial queries, Google's own ads inside the box), a visibility channel with worse click economics than the one it replaced. Being cited still puts your brand inside an answer surface Google says reaches [over 2.5 billion users a month](https://blog.google/products-and-platforms/products/search/new-controls-website-owners/), still reaches the minority who do click through, and increasingly feeds the same retrieval systems behind AI Mode and other engines. ## Can You Measure Your AI Overview Visibility? Partly, and the "partly" changed in June 2026, so most advice you will read is out of date in one direction or the other. For two years the answer was simply no: Google Search Console counted AI Overview appearances inside the standard Performance report, folded into ordinary web results with no way to isolate them. Then in June 2026, Google [introduced generative-AI performance reports](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports) in Search Console, rolling out gradually. They show impressions for AI Overviews and AI Mode, counted together, broken down by page, country, and device. An impression there means a URL from your site actually appeared inside an AI feature, per [Google's definition](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports). What the report still does not show: clicks, clickthrough rate, position, which answer your page appeared in, what that answer said, or who was cited alongside or instead of you. You can see that your pages surfaced somewhere in Google's AI features; you cannot see the answers themselves or the competitors who own the rest of them. Closing that gap takes direct observation: running your target queries, recording whether an overview fires, and logging which domains it cites over time, since the answers themselves are volatile and personalized. That is scriptable by hand, and it is what [AI Overview tracking tools](https://geotoolbox.ai/blog/ai-overview-tracker) automate. It is also the gap geotoolbox exists for: the failure mode where an AI Overview cites a competitor above your #1 ranking is invisible in every report Google gives you, so it has to be observed from the outside. The wider playbook, from picking prompts to baselining share of voice, is in our guide to [tracking brand mentions in AI search](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search). ## AI Overviews vs AI Mode, Featured Snippets, and ChatGPT "Is this just ChatGPT in Google?" is one of the most common questions about AI Overviews (it shows up in Google's own People Also Ask box), and the answer is no in every row of this table:
SurfaceWhat it isWhere answers come fromConversational?
AI OverviewGenerated summary on top of the classic results pageGemini grounded in live Google Search retrieval, with citationsNo, one-shot per query
AI ModeSeparate conversational search tabSame Gemini-plus-Search stack, heavier query fan-outYes, keeps context across turns
Featured snippetVerbatim excerpt box from a single pageOne ranked page, quoted, not generatedNo
ChatGPTStandalone chatbot with optional web searchIts own model and retrieval, no Google indexYes
The line between the first two is thinning. A "Show more" tap in an AI Overview now drops you into AI Mode with your query carried over, and Search Console already reports the two surfaces as one bucket. They remain distinct surfaces with only partially overlapping citations, which is why visibility in one does not guarantee visibility in the other. The featured snippet comparison matters for a different reason: a snippet lifts your words and attributes a single source, while an AI Overview paraphrases you into a blended answer alongside several other sources. Same real estate, much weaker claim on the click. ## Frequently Asked Questions ### How do you turn off Google AI Overviews? You mostly cannot. There is no account setting that disables AI Overviews globally; workarounds include Google's "Web" results filter, adding `udm=14` to search URLs, and browser extensions that hide the box. None of them is a full off switch; our guide to [turning off AI Overviews](https://geotoolbox.ai/blog/how-to-turn-off-ai-overviews) covers what each option does and does not do. ### What triggers a Google AI Overview to appear? Google shows one when its systems judge a generated answer more useful than links alone. In practice that means informational, question-form, longer, and more complex queries trigger them most, while navigational, local, news, and shopping queries usually do not. ### Should you trust what an AI Overview says? Use it for orientation, verify it for decisions. The text is generated and grounded in retrieved pages, so it is usually right, but independent tests keep finding meaningful error rates on specialist topics, and the sources are one tap away. ### Is an AI Overview the same as ChatGPT? No. ChatGPT is a standalone conversational model with its own retrieval; an AI Overview is a one-shot summary generated by Gemini from live Google Search results and pinned above the classic listings, with citations to the pages it drew on. ### Why did AI Overviews stop showing for you? Trigger thresholds shift constantly, and results vary with query phrasing, language, and location. A query that fired an overview last month may return plain results today; that volatility is normal and is one reason tracking them requires repeated sampling. ### Do AI Overviews use your site's content without permission? If your pages are indexed for Google Search, they are eligible for AI Overviews; there is no AIO-specific opt-out that preserves your rankings, though Google began testing one with UK site owners in June 2026 and it has since spread to some accounts outside the UK too, with the rollout still partial. The blunt instruments cost visibility either way: `noindex` removes the page from Search entirely, while `nosnippet` keeps it indexed but suppresses its snippet and its use in AI features. That trade-off is what publishers are litigating. ## The Answer Box Is the New Front Page So, what are Google AI Overviews? The generated layer Google now puts between searchers and your site: Gemini summaries grounded in regular search retrieval, triggered disproportionately on informational, longer, higher-volume queries, cited from pages that do not always match the ranking order. They are not going away, and nothing on the current trajectory reverts the click economics. The lever you have is being the source the answer quotes, and knowing when it quotes someone else instead. If you want to see where your site stands today, run it through geotoolbox's free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness); it takes a minute and checks five reachability foundations that decide whether AI search engines can read your pages at all. ## Sources - Google AI Overviews: Search anything, effortlessly - Google - `search.google/ways-to-search/ai-overviews` - AI Overviews: About last week - Elizabeth Reid, Google, May 2024 - `blog.google/products-and-platforms/products/search/ai-overviews-update-may-2024` - Google users are less likely to click on links when an AI summary appears in the results - Pew Research Center, July 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` - Update: AI Overviews Reduce Clicks by 58% - Ahrefs, February 2026 - `ahrefs.com/blog/ai-overviews-reduce-clicks-update` - Insights from 56 Million AI Overviews - Ahrefs, May 2025 - `ahrefs.com/blog/insights-from-56-million-ai-overviews` - Semrush AI Overviews Study - Semrush, December 2025 - `semrush.com/blog/semrush-ai-overviews-study` - Google Restricts AI Search Tool After Nonsensical Answers - Forbes, May 2024 - `forbes.com/sites/roberthart/2024/05/31/google-restricts-ai-search-tool-after-nonsensical-answers-told-people-to-eat-rocks-and-put-glue-on-pizza` - Rolling Stone owner Penske Media sues Google over AI summaries - TechCrunch, September 2025 - `techcrunch.com/2025/09/14/rolling-stone-owner-penske-media-sues-google-over-ai-summaries` - Chegg sues Google for hurting traffic as it considers alternatives - CNBC, February 2025 - `cnbc.com/2025/02/24/chegg-sues-google-for-hurting-traffic-as-it-considers-alternatives.html` - Google faces EU antitrust complaint over AI Overviews - Search Engine Land, July 2025 - `searchengineland.com/google-faces-eu-antitrust-complaint-over-ai-overviews-458123` - Google strikes $60m deal with Reddit for AI training data - Tom's Guide, February 2024 - `tomsguide.com/ai/google-strikes-dollar60m-deal-with-reddit-for-ai-training-data-what-you-need-to-know` - AI Overviews: Everything You Need to Know - WordStream, April 2026 - `wordstream.com/blog/google-ai-overviews` - Introducing Search Generative AI performance reports in Search Console - Google Search Central, June 2026 - `developers.google.com/search/blog/2026/06/gen-ai-performance-reports` - Google Search's I/O 2026 updates - Google, May 2026 - `blog.google/products-and-platforms/products/search/search-io-2026` - New opportunities, control and insights for website owners - Google, June 2026 - `blog.google/products-and-platforms/products/search/new-controls-website-owners` - AI Overviews - Wikipedia - `en.wikipedia.org/wiki/AI_Overviews` --- ## What Is Google AI Mode? How It Works > Google AI Mode is Google's conversational, Gemini-powered search. How query fan-out works, how it differs from AI Overviews, and how to turn it on or off. - Canonical: https://geotoolbox.ai/blog/what-is-google-ai-mode - Published: 2026-07-23 · Updated: 2026-08-05 Google AI Mode is the biggest change to how Google Search works in years, and most people are not sure what it actually is. This is a plain explainer: what Google AI Mode is, how it works under the hood, how it differs from AI Overviews and the Gemini app, and how to turn it on (or get back to plain links). It is mostly about what the thing is, with a short note at the end on what it means if you run a website. ## What Google AI Mode Is Google AI Mode is a conversational, AI-first way to search, built into Google Search. Instead of returning a page of blue links, it uses Google's Gemini models to read across the web and hand you a written answer with a few cited sources, and you can keep asking follow-up questions like a chat. Google calls it its "most powerful AI search experience." AI Mode is still Google Search, not a separate chatbot. It runs on Google's live search index, its ranking systems, the Knowledge Graph, and real-time data like shopping and local results. The difference is the interface: you get a synthesized answer and a conversation rather than a page of links to click through yourself. You can ask by typing or talking, and you can add a photo, an image, or a document and ask about it. You reach it as its own tab or page. When you open AI Mode and ask something, the AI response becomes the whole results experience, with links moved into and beneath the answer rather than listed down the page. It is designed for the harder questions, the ones with several parts, comparisons, or trade-offs that used to take five separate searches and a dozen open tabs. ## How Google AI Mode Works: Query Fan-Out The mechanism behind AI Mode is a technique Google calls query fan-out. When you ask a question, Google does not run one search. It breaks your question into many smaller sub-questions and, [in Google's own words](https://blog.google/products-and-platforms/products/search/ai-mode-search/), issues "multiple related searches concurrently across subtopics and multiple data sources and then brings those results together." A single prompt can trigger many searches running at once. Here is the sequence. Your question goes in. AI Mode decomposes it into related sub-queries and fires them in parallel against Google's index and live data. It pulls back passages from across many pages, then a Gemini model reasons over everything it gathered and writes one answer, attaching the sources it leaned on. Ask a follow-up and the whole fan-out runs again, this time with the context of the conversation so far.
Diagram of Google AI Mode query fan-out: one question splits into parallel sub-searches that Gemini synthesizes into one cited answer
Query fan-out: one question becomes many parallel searches, then Gemini synthesizes a single cited answer.
Two things follow from this design. First, it reaches deeper into the web than a normal search, surfacing pages that would never have ranked on the first results page for your exact wording. Second, because it is stitching together passages from many sources, AI Mode handles longer, messier questions well, the kind that run much longer than a typical keyword search. Google has not published how many sub-queries a fan-out generates, and the number varies with the question. If you want the mechanics in more depth, see our explainer on [query fan-out](https://geotoolbox.ai/blog/query-fan-out). ## AI Mode vs AI Overviews vs Gemini vs Classic Search Most of the confusion about AI Mode comes from Google using several similar-sounding names for different things. There are really four surfaces, and they are not the same product. **AI Overviews** is the short AI summary that appears at the top of normal search results for some queries. You did not ask for it, it just shows up above the links, and it is mostly a one-shot answer. **AI Mode** is separate: a dedicated tab or page you deliberately open, where the AI answer is the entire experience and you can keep the conversation going. **Gemini** is two things at once, which does not help: it is Google's standalone chatbot app, and it is also the name of the model family that powers both AI Overviews and AI Mode. **Classic Search** is the familiar ranked list of blue links.
SurfaceHow you get itWhat you seeConversation?
Classic SearchDefault results pageRanked list of linksNo, one query at a time
AI OverviewsAuto-appears atop some resultsA short AI summary above the linksMostly one-shot
AI ModeYou open the AI Mode tab or pageA full AI answer with cited linksYes, multi-turn follow-ups
Gemini appSeparate chatbot at gemini.google.comA chat assistant, not a search pageYes, but it is not Search
The simplest way to keep them straight: AI Overviews comes to you, AI Mode is somewhere you go, and Gemini is both a separate app and the engine under the hood. If you mainly want to know how AI Overviews behaves, we cover that in [what Google AI Overviews are](https://geotoolbox.ai/blog/what-are-google-ai-overviews), and the model itself in our guide to [what Gemini is](https://geotoolbox.ai/blog/what-is-gemini). ## How AI Mode Evolved: From SGE to Gemini 3 If you remember an earlier "Search Generative Experience," that is where this started. SGE was Google's opt-in Labs prototype that generated summaries over search results. Google retired the SGE name and shipped the summary part broadly as AI Overviews. AI Mode is the newer, full conversational surface that came after. So AI Mode is not the same as SGE or the old Bard, it is what that line of experiments grew into. The model powering it has moved fast, which is why other explainers disagree on the version. Google [launched AI Mode on March 5, 2025](https://blog.google/products-and-platforms/products/search/ai-mode-search/) as an opt-in Search Labs experiment in the US, running on a custom version of Gemini 2.0. Around Google I/O in May 2025 it [began rolling out to all US users](https://blog.google/products/search/google-search-ai-mode-update/), no Labs sign-up required, and was upgraded to a custom Gemini 2.5 model. Google's [product page](https://search.google/ways-to-search/ai-mode/) still describes it as using "Gemini 3's next-generation intelligence," but that page has fallen behind: at [I/O 2026](https://blog.google/products-and-platforms/products/search/search-io-2026/) Google made **Gemini 3.5 Flash** the default model in AI Mode for everyone globally. So, which Gemini model is it? 2.0 at launch, 2.5 by mid-2025, Gemini 3 through late 2025, and Gemini 3.5 Flash as the default since May 2026, with Gemini 3 Pro available as a paid opt-in from the model menu.
WhenWhat happenedModel
2023-2024Search Generative Experience (SGE) runs as a Labs prototypePaLM 2, then Gemini
Mar 5, 2025AI Mode opens as a US Search Labs opt-in experimentCustom Gemini 2.0
May 2025 (I/O)Broad US rollout begins, no Labs sign-up neededCustom Gemini 2.5
Aug 2025Expands to 180+ countries and territories, in EnglishGemini 2.5
2026AI Overviews and AI Mode integrate; passes one billion monthly usersGemini 3.5 Flash (default)
The scale grew just as fast. Google said AI Mode passed [one billion monthly users](https://blog.google/products-and-platforms/products/search/search-io-2026/) by 2026, while AI Overviews, the surface built into normal results, reaches an even larger audience. In under two years this went from a gated Labs toggle to something a billion people use every month. At the same I/O, Google tied the two surfaces together: you can now ask a follow-up straight from an AI Overview and continue in AI Mode, so they increasingly feed one flow even though they remain distinct entry points. ## How to Turn On (Access) Google AI Mode In markets where it has rolled out, AI Mode is less a setting you flip than a place you open. There are three common ways in: 1. The **AI Mode tab** at the top of a normal Google results page, usually next to "All" and the other filters 2. Going straight to **google.com/ai** 3. The AI Mode entry point in the **Google app** on your phone If you do not see it yet, the reason is almost always availability rather than a hidden setting. AI Mode started in the US, expanded to [180-plus countries in English](https://searchengineland.com/google-launches-ai-mode-in-180-countries-and-territories-461040) in August 2025, went global in Spanish that September, and now covers more than 190 countries and territories in 100-plus languages. It also has not come to Workspace (work and school) Google accounts at the same pace as personal ones, so a signed-in work account may not show it even where it is live. Where a feature is still rolling out, Search Labs, Google's opt-in area for experiments, is often where it appears first, so checking there can surface newer AI Mode capabilities early. ## Is Google AI Mode Free? Yes. AI Mode is free to use as part of Google Search. You do not need a subscription to open the tab and ask questions. There are paid tiers, but they buy extras rather than access. Google AI Plus, Google AI Pro, and Google AI Ultra add things like Deep Search, which runs a longer, report-style research pass over many more searches, plus the option to run AI Mode on Gemini 3 Pro, higher usage limits, and some of the newer agentic features (such as AI Mode carrying out tasks like booking) that have been rolling out to subscribers first. The Gemini 3 Pro option is region-gated (much of the Americas has it; no EU country does) and English-only for now. The core experience, asking a question and getting a synthesized, cited answer, costs nothing. ## Can You Turn Off Google AI Mode? This is the most common frustration, and the honest answer is that there is no single official switch to permanently disable AI Mode. Google has not shipped a "turn off AI Mode" toggle in search settings. You rarely need one, though, because AI Mode is a separate tab rather than something forced into your normal results. If you want plain links, just stay on the "All" tab and ignore the AI Mode tab, or use the **Web** filter, which strips a result page down to classic links only. Power users reach the same view by adding **&udm=14** to a Google search URL. A few browser-level tricks (a Chrome flag, a Search Labs toggle) can also hide AI features, but they are experimental and tend not to stick across updates. None of this makes the AI Mode tab disappear, but it keeps your day-to-day searching on plain, ranked links. Worth separating out: the AI summary that appears at the top of ordinary results is AI Overviews, not AI Mode, and getting rid of that is a different question with its own partial workarounds. We cover those in detail in [how to turn off AI Overviews](https://geotoolbox.ai/blog/how-to-turn-off-ai-overviews). ## Is Google AI Mode Accurate? What It Gets Wrong Google itself is upfront here: its [AI Mode page](https://search.google/ways-to-search/ai-mode/) states plainly that "AI Mode is experimental and may make mistakes." Treat it as a fast, well-sourced starting point, not a final authority. The known weak spots are the ones common to all generative search. It can hallucinate, stating something false or outdated with full confidence, which matters most on your-money-or-your-life topics like health, finance, and legal questions. Its citations are approximate: because fan-out pulls passages from many pages, the links attached to an answer do not always map cleanly to the specific sentence they supposedly support. And like any generative system, it can read as more opinionated or more certain than the underlying evidence warrants. The counterweight is that AI Mode is grounded in Google's live index rather than working from training data alone, which tends to make it better sourced than a standalone chatbot. It is not a settled experience, though: [Nielsen Norman Group's usability review](https://www.nngroup.com/articles/google-ai-mode/) found the search itself powerful but the interface uneven, so the answer is often stronger than the way you have to work with it. The practical habit is the same one good researchers already use: read the answer, then click the cited links to confirm anything that matters. ## What AI Mode Means for Being Found If you run a website, the shift to grasp is that AI Mode answers questions with citations, not a ranked list, so the thing you are competing for is being one of the sources it pulls into the answer. That is a different job from ranking first, and ranking first does not guarantee it. It starts with something more basic than tactics: your pages have to be reachable by Google's crawlers at all. A common failure we see when scanning sites for AI visibility at geotoolbox is not weak content but a robots or crawler rule quietly blocking the bots that feed these systems. That is worth checking before anything else. This is only the "what it is" side of the story. The actual playbook for getting cited and measuring it lives in our guide to [Google AI Mode SEO](https://geotoolbox.ai/blog/google-ai-mode-seo). ## The Short Version Google AI Mode is Google's conversational search: a Gemini-powered tab that fans your question out into many parallel searches and hands back one cited answer you can keep questioning. It is free, it is separate from both AI Overviews and the Gemini app, it has run on Gemini 3.5 Flash as its default model since May 2026, and it is experimental enough that you should still click the sources. If you would rather have plain links, the "All" or Web view is always a tab away. One thing you can do today is make sure AI Mode can actually see your site. Run your domain through our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) to confirm the bots behind Google's AI answers are not being blocked. ## Frequently Asked Questions ### Is Google AI Mode the same as Gemini? No, though they are related. Gemini is Google's standalone chatbot app, and it is also the name of the model family that powers AI Mode. AI Mode is a search experience inside Google Search that uses a Gemini model, combined with Google's live index and ranking, to answer queries with cited links. The Gemini app is a separate product you visit at gemini.google.com. ### Is Google AI Mode free? Yes. AI Mode is free to use as part of Google Search, with no subscription required. Paid Google AI Plus, AI Pro, and AI Ultra plans add extras like Deep Search, more capable models, higher limits, and some agentic features, but the core question-and-answer experience costs nothing. ### Can I turn off Google AI Mode? There is no official setting that permanently disables it. Because AI Mode is a separate tab, the practical fix is to stay on the "All" tab or use the Web filter for classic links, or add &udm=14 to a search URL. Turning off the AI Overviews summary at the top of normal results is a different question, covered in our [guide to turning off AI Overviews](https://geotoolbox.ai/blog/how-to-turn-off-ai-overviews). ### Is Google AI Mode replacing normal Google Search? Not entirely, at least not yet. AI Mode is an additional way to search that sits alongside classic results, and the ranked-link view is still there whenever you want it. That said, Google keeps folding AI deeper into the default experience: at I/O 2026 it connected AI Overviews and AI Mode so a follow-up from an Overview flows into an AI Mode conversation, so the surfaces are converging even as the plain-links view stays available. ### Does Google AI Mode use my personal data? Standard AI Mode answers from public web results, not your private data. Google has added an opt-in Personal Intelligence layer that can connect Gmail, Photos, and other Google apps to personalize answers, but it is off until you turn it on, and you can manage or disable it in your Search personalization settings. ### Is Google AI Mode accurate? It is usually well sourced because it grounds answers in Google's live index, but Google itself labels it experimental and warns it "may make mistakes." It can hallucinate, its citations are approximate, and it can sound more certain than it should. Verify anything important by clicking the cited links, especially for health, finance, or legal questions. ### What countries and languages is Google AI Mode available in? It launched in the US, expanded to more than 180 countries in English in August 2025, went global in Spanish in September 2025, and now covers more than 190 countries and territories in 100-plus languages. Availability also depends on account type, and it has reached personal Google accounts ahead of Workspace (work and school) accounts. ## Sources - Introducing AI Mode in Search - Google, March 5, 2025 - `blog.google/products-and-platforms/products/search/ai-mode-search/` - Google AI Mode product page (current capabilities and model) - Google - `search.google/ways-to-search/ai-mode/` - What's new in Search at I/O 2026 (usage scale) - Google - `blog.google/products-and-platforms/products/search/search-io-2026/` - Get AI-powered responses with AI Mode in Google Search - Google Search Help - `support.google.com/websearch/answer/16011537` - AI Mode in Search update (I/O 2025 rollout and Gemini 2.5) - Google, May 20, 2025 - `blog.google/products/search/google-search-ai-mode-update/` - Google launches AI Mode in 180 countries and territories - Search Engine Land - `searchengineland.com/google-launches-ai-mode-in-180-countries-and-territories-461040` - Google AI Mode: Powerful Search, Poor Usability - Nielsen Norman Group - `nngroup.com/articles/google-ai-mode/` --- ## AI Brand Sentiment: How to Track What AI Says > What AI brand sentiment is, why social listening misses it, and a free method to track and improve how ChatGPT, Gemini, and Perplexity describe your brand. - Canonical: https://geotoolbox.ai/blog/ai-brand-sentiment - Published: 2026-07-22 · Updated: 2026-08-01 AI assistants do not just mention your brand. They frame it: as the safe choice, the budget option, the one with the learning curve, or the one they quietly leave out. AI brand sentiment is the measure of that framing, and unlike a ranking, you can read it, score it, and change it. Your first baseline takes an afternoon, not a platform. The whole method in one line: run 20 to 40 fixed buyer prompts across at least 3 engines, 3 times each, classify every brand mention on a five-level scale, and score net sentiment as (recommended + positive − negative) ÷ total mentions × 100, tracked monthly. ## What Is AI Brand Sentiment? **AI brand sentiment** is the tone an AI assistant uses when it talks about your brand: whether ChatGPT, Gemini, Perplexity, or Claude describes you as a recommendation, an option, or a warning. It is brand perception, as synthesized by the models your buyers ask. It is not the same thing as showing up. Mention rate tells you whether your brand appears in AI answers at all. Sentiment tells you what those appearances are doing for you. A brand that shows up in most category prompts looks visible until you read the answers and find most of them hedge, caveat, or recommend someone else. The distribution matters here. A [February 2026 analysis of 1.8 million brand-mentioning AI responses](https://rocketblue.ai/articles/tracking-brand-mentions-in-ai-chatbots-a-comprehensive-guide-to-monitoring-brand-presence-in-chatgpt-responses-feb-2026-data/) by rocketblue found that 80.6% of brand mentions in AI answers are neutral, 18.4% positive, and about 1% negative. Openly negative framing is rare. The real fight is moving your brand out of the neutral pile and into the positive tail, where the engine frames you favorably or recommends you outright. Most AI sentiment analysis tools collapse this into positive, neutral, and negative. A five-level scale keeps the resolution where the actionable information lives:
LevelSignal languageWhat it means
Recommended"the best choice for", "widely recommended", "trusted by"The engine endorses you for the use case
Positive framing"strong at", "known for", "a solid option"Favorable attributes, no explicit recommendation
Neutral listing"options include A, B, and C"Mentioned without evaluative framing; undifferentiated in the answer
Hedged"may be suitable for", "some users prefer", "worth considering but"The answer qualifies your suitability; often a weak-signal symptom
Negative"lacks", "users report issues with", "not recommended for"Active warning; buyers deprioritize you before ever visiting your site
The tie-breaker between the top two levels: **Recommended** requires an explicit recommendation aimed at a use case; praise without one is Positive framing. One thing this scale deliberately excludes: factual errors, the hallucinations that get your pricing wrong or describe a discontinued product as current. Those are an **accuracy** problem, not a sentiment level. Track wrong claims as their own count; the fix is corrections at the source, not positioning work. New to measuring this? Start with [what AI visibility is](https://geotoolbox.ai/blog/what-is-ai-visibility); sentiment is the quality layer on top. ## Why Social Listening Tools Miss It Your social listening stack does not cover this. Social media monitoring platforms like Brandwatch, Sprout Social, and Hootsuite track consumer sentiment: what **people** say about your brand on social platforms, forums, and review sites. AI brand sentiment is what the **models themselves** say when a buyer asks them a question; no volume of social monitoring surfaces it. The two signals differ in every way that matters for measurement:
DimensionSocial listeningAI brand sentiment
Signal sourceHuman posts, reviews, customer feedbackSynthesized answers from ChatGPT, Gemini, Perplexity, Claude
VolumeThousands of mentions a dayOne answer per prompt, per engine, per run
How it changesIn real time, with human activityIn steps: model updates, retraining, and shifts in the sources engines retrieve
The fixReply, moderate, manage brand reputationChange what the engines read: owned content, third-party consensus, corrections
When it hurts youAfter a post spreadsBefore the buyer ever reaches your site
The diagnosis logic differs too. A brand with warm social sentiment and hedged AI sentiment does not have a customer satisfaction problem. It usually has a content and entity-consistency problem: the engines are not finding confident, corroborated claims about it. Read the answers and their cited sources before deciding which. Social sentiment analysis stays relevant for what humans say; it simply cannot see what the models say. The two need different work, owned by different teams, on different cadences. ## How AI Assistants Form an Opinion of Your Brand The stakes first. In a [Semrush survey of 1,030 US shoppers who had tried AI tools](https://www.semrush.com/blog/ai-tools-the-modern-buyer-journey-study/), run in December 2025, 57% used AI to narrow down their choices, 53% to compare products they were already considering, and 50% to make a final decision. [BCG's 2026 consumer research](https://www.bcg.com/publications/2026/consumers-trust-ai-to-buy-better-brands-must-adapt) adds that shopping-related GenAI use grew 35% between February and November 2025, and more than 60% of consumers express high trust in what the tools tell them. That opinion is assembled from four main inputs: training data, live retrieval, structured data on your pages, and the third-party consensus (reviews, comparisons, forums, press) that shapes your brand image. Each engine weighs them differently, so the same prompt produces different sentiment per platform.
EngineLeans onWhat that means for your sentiment
ChatGPTTraining data, plus web search on demandCan carry a cached, outdated impression of you; fixes lag until it searches or retrains
Gemini and Google AI OverviewsGoogle's search index and Knowledge Graph (separate products, shared grounding)Your entity consistency across Google surfaces shapes the framing
PerplexityRetrieval-led answers with visible citationsOften first to reflect new content, and heavily dependent on what its citations say about you
ClaudeTraining data, plus web searchFraming tends to track its training corpus when it does not search
The variance starts with whether brands get named at all. Across rocketblue's tracked prompt set in February 2026, [Claude named brands in 97.3% of responses while AI Overviews did in 48.5%](https://rocketblue.ai/articles/tracking-brand-mentions-in-ai-chatbots-a-comprehensive-guide-to-monitoring-brand-presence-in-chatgpt-responses-feb-2026-data/), with ChatGPT at 73.6%. An engine that names brands half as often also gives you half the sample, so each engine needs its own reading. A sentiment reading from one engine should never be assumed to generalize to the others, so any serious [AI visibility tracking](https://geotoolbox.ai/blog/how-to-track-ai-visibility) has to be multi-engine. And the same engine will not give the same answer twice: sampling variation means sentiment flickers between runs. Both are method problems, solved next. ## How to Track AI Brand Sentiment Step by Step A trustworthy sentiment baseline needs a fixed prompt set, a repetition rule, a classification rubric, and a spreadsheet. No platform required. ### 1. Build a Prompt Set from Real Buyer Questions Write 20 to 40 prompts that mirror how buyers ask. Pull them from four places: high-intent Google Search Console queries rephrased as questions, sales-call questions, support tickets, and direct reputation probes. Cover four types: category prompts ("best [category] tools for [use case]"), comparison prompts ("[you] vs [competitor]"), reputation prompts ("why do people switch away from [brand]"), and negative probes ("which [category] tools should I avoid", "which [category] tools are overpriced for what they deliver"). The negative probes matter most; they surface associations the polite prompts never show. [Query fan-out](https://geotoolbox.ai/blog/query-fan-out) shows how to expand the set the way engines themselves do. Then freeze the set. A prompt list you rewrite every month cannot show you a trend. ### 2. Run Every Prompt on at Least 3 Engines, 3 Times Each Run the set on ChatGPT, Gemini, and Perplexity at minimum, in fresh logged-out sessions so personalization does not contaminate the reading. The part almost everyone skips: run each prompt **three times per engine**. LLMs sample; a single run is an anecdote, not a measurement. One run out of three where your brand drops from an answer is likely noise. The same drop across all three runs, or across two engines, is worth treating as signal. Three runs is a floor, not statistics: it filters the coin-flip flicker, and trends across monthly cycles do the rest. That one rule separates a sentiment tracker from a mood ring, and the vendors' own walkthroughs rarely mention it. Capture full responses with metadata: prompt, engine, date, and which sources the answer cited. The score you compute next is only useful if you can go back and read why it moved. ### 3. Classify Every Mention Label each brand mention with the five-level scale from earlier: recommended, positive, neutral, hedged, negative. Flag factually wrong claims separately as accuracy errors. You can use an LLM as the classifier; sentiment analysis of short text against a fixed rubric is exactly the job it is good at. Paste the rubric and the response, ask for a label, and spot-check 10% by hand. A curiosity for the technically minded: in LLaMA-family models, a [2025 probing study](https://arxiv.org/abs/2505.16491) found sentiment signals read directly from hidden layers beat prompting-based classification by up to 14%. For this workflow, a rubric plus human spot-checks carries the load; keep the classifier consistent: same model, same rubric, every cycle. ### 4. Score It with One Consistent Formula Use one formula and never change it: **Net sentiment = (recommended + positive − negative) ÷ total brand mentions × 100** Neutral and hedged mentions count in the denominator but score zero: they signal absent conviction rather than damage. Track accuracy errors as a separate rate (wrong claims ÷ total mentions), not blended into sentiment. A worked example, using a distribution close to the real-world baseline above: 120 classified mentions, of which 6 recommended, 16 positive, 82 neutral, 13 hedged, 3 negative. Net sentiment = (6 + 16 − 3) ÷ 120 × 100 = **+16**, rounded from 15.8.
Net sentimentReading
+40 and aboveStrongly favorable framing; protect the sources doing the work
+15 to +39Net positive, with a large convertible neutral pool
−15 to +14Undifferentiated: engines know you exist but have nothing to say
Below −15Active negative narratives; find the sources feeding them before writing anything
The bands are directional: there is no industry-standard sentiment score, which is why your formula staying constant matters more than which one you pick. Read any band next to the full category counts and the sample size, never alone.
![Pipeline for tracking AI brand sentiment: prompt set, multi-engine repeated runs, mention classification, net sentiment score, driver analysis.](/blog/ai-brand-sentiment/ai-brand-sentiment-scoring-pipeline.png)
Five stages from prompt set to owned fixes: the repetition in stage two is what makes the score in stage four trustworthy.
### 5. Check the Sentiment of What Engines Cite Read the sources your captured answers cited and note whether each one frames you positively or negatively. This is the layer most tracking misses: an engine that keeps retrieving a lukewarm comparison post keeps producing lukewarm answers about you. Your fix list starts from these pages: treat each one as a lead to verify, not a proven cause. The mechanics of finding them are covered in our [brand mention tracking guide](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search). ### 6. Set a Cadence and a Baseline Score monthly; bi-weekly while actively repairing something, and re-run within a week of major model releases. If Google Search Console has rolled its [generative AI performance report](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports) out to your property, wire it in: it shows AI Overviews and AI Mode impressions for free, revealing which pages the AI layer surfaces. ## Where Sentiment Fits in Your AI Visibility Stack Sentiment is the fourth metric of four, the one that gives the other three their meaning. **Mention rate** answers: do engines know we exist? It is the brand awareness layer. **[AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice)** answers: how often do they surface us versus competitors? **Citation rate** answers: do our pages get used as sources? **Sentiment** answers the question the others can't: when we do show up, is it helping? (Position within the answer matters too; treat prominence as part of mention tracking, not sentiment.) The metrics interact. High [share of voice](https://geotoolbox.ai/glossary/share-of-voice) with hedged sentiment is worse than modest visibility with confident recommendations: you are spending exposure teaching buyers to be unsure about you. Low mention rate with strong sentiment usually means the fix is distribution rather than positioning. And rising citations with flat sentiment points at how your cited pages describe you. Read them together, in order: existence, frequency, attribution, quality. Alone, a sentiment score is trivia; next to the other three, it is a diagnosis. ## How to Fix Negative AI Brand Sentiment ### Find the Driver Before Writing Anything Group your hedged and negative mentions by subject: features, pricing, support, ease of use, or trust. Each cluster has a different owner and a different fix: recurring pricing complaints belong to marketing, feature gaps to product, "slow support" themes to customer success. The grouping usually collapses the problem into two or three narratives with names attached. Then trace each narrative to its sources. Your captured responses include citations; read them. A recurring "steep learning curve" theme often lives in a handful of reviews and comparison posts the engines keep retrieving. Start with the review profiles your answers cite most (for B2B software, usually G2 and Capterra), then the two or three comparison posts that keep reappearing. Those URLs are your work list. ### What Moves the Needle The levers: **Precision in owned content.** Vague owned pages give engines nothing confident to retrieve. Replace marketing abstractions with checkable specifics: who the product is for, what it does and does not do, current pricing, named capabilities. **Third-party consensus.** Engines trust corroboration more than self-description. Current profiles on the review sites your captured answers cite, coverage in publications the engines retrieve, and accurate comparison content typically do more for sentiment than any page on your own domain; [getting cited by AI](https://geotoolbox.ai/blog/how-to-get-cited-by-ai) is its own playbook. This is slow, and for most brands it is the main lever. **Correcting wrong claims at the source.** For accuracy errors, find the cited page carrying the wrong fact and send its author a short note quoting the wrong sentence and linking the current fact, then fix every inconsistency on your own properties that could have seeded it. Accuracy work compounds: buyers distrust AI answers already, and per [Gartner's May 2026 B2B buyer survey](https://martech.org/b2b-buyers-trust-ai-less-than-marketers-think/), more than half of B2B buyers say they are more likely to encounter misleading information from AI tools than from a sales rep, and 69% turn to a sales rep to validate AI-generated insights. ### What to Expect Timelines depend on the engine's plumbing. Retrieval-led engines like Perplexity can reflect new content within weeks. Narratives answered from training data tend to persist until a model update, which you do not control and cannot schedule. Consensus problems take months; budget accordingly. What does not work: chasing single-run fluctuations (if a change does not survive your repetition rule, it is not real), and the lone rebuttal page (sentiment follows the weight of consensus across sources). ## AI Brand Sentiment Tools The tool landscape sorts into three tiers, and two of them get conflated constantly. **AI-native trackers** run prompt sets across answer engines on a schedule and classify the LLM mentions with natural language processing (NLP): Profound, Peec AI, Otterly, and a long tail of newer entrants. **Answer engine optimization (AEO) add-ons** bolt AI answer monitoring onto an existing SEO suite, the way [Semrush and Ahrefs](https://geotoolbox.ai/blog/semrush-vs-ahrefs) have. **Social listening platforms** are the tier that does not belong in this conversation: Brandwatch and Sprout measure human posts, not model answers. How fragmented is the category? When we put the same sentiment-tracking questions to ChatGPT, Gemini, Perplexity, and Claude in July 2026 while researching this article, no tool was named by all four engines, and any two engines' primary tool lists overlapped by at most a couple of names. The engines have not settled on who does this well; weight any vendor's "leading platform" claim accordingly. We keep a maintained comparison in our [Profound alternatives](https://geotoolbox.ai/blog/profound-alternatives) breakdown. Full disclosure on where we sit: geotoolbox tracks brand mentions and share of voice across eight AI engines, keeping the verbatim phrasing behind every mention. We do not sell a sentiment score; that is why this article hands you the rubric and formula: run them over your captured answers, or over the phrasing we store, and the number is yours to audit. Whichever tier you pick, apply the method test: does it store full responses, not just scores? Does it run prompts more than once? Can you export the raw data? Hold us to it too: geotoolbox keeps verbatim phrasing rather than complete responses, enough to audit a mention but less than a full transcript. A sentiment score you cannot audit back to the answers that produced it belongs in the vendor's pitch deck. ## Frequently Asked Questions ### What is a good AI brand sentiment score? On the net sentiment formula in this article, +40 or above means engines actively recommend you, +15 to +39 is net positive with room to convert neutral mentions, and anything below −15 signals active negative narratives. There is no industry-standard score and tools compute on incompatible scales, so trend against your own baseline. ### Can ChatGPT do sentiment analysis? Yes, and it is a reasonable classifier for this workflow: give it your rubric and a captured response, ask for a label, and spot-check a sample by hand. Use the same model and rubric every cycle so the trend stays comparable, and never ask an engine to assess its own opinion of your brand; classify captured responses instead. ### Which AI engine should you track first? The one your buyers use: for most brands ChatGPT first, then Google's AI surfaces, then Perplexity. But single-engine tracking misleads: engines name brands at rates from 97.3% (Claude) down to 48.5% (AI Overviews), and their sentiment toward the same brand differs. Three engines is the practical floor. ### How often should you track AI brand sentiment? Monthly as a baseline, bi-weekly while repairing a narrative, and within a week of any major model release. Absent a model release or a news event, more frequent checking mostly measures sampling noise unless you also increase runs per prompt. ### How long does it take to change what AI says about your brand? Often weeks on retrieval-led engines like Perplexity, where corrected content shows up once crawled and cited. Months on narratives ChatGPT or Claude answer from training data, which tend to persist until a model update you cannot schedule. Accuracy corrections move fastest, consensus problems slowest. ## Start with 20 Prompts and a Spreadsheet The method above scales down: 20 prompts, 3 engines, 3 runs, one spreadsheet, with an LLM classifier doing the labeling. Running it is an afternoon; keeping it current every month is the part worth automating. That first baseline tells you whether you are unknown, undifferentiated, or actively framed; everything after is trend lines and driver work. geotoolbox's [domain overview](https://geotoolbox.ai/features/domain-overview) handles that running layer: it scans eight AI engines for your brand's mentions and share of voice, keeping the verbatim phrasing behind each mention, so reading why a number moved stays a click away. The sentiment layer is the formula from this article, applied on top. ## Sources - How AI Tools Influence the Modern Buyer Journey - Semrush, December 2025 - `semrush.com/blog/ai-tools-the-modern-buyer-journey-study/` - Consumers Trust AI to Buy Better. Brands Need to Move Quickly. - BCG, 2026 - `bcg.com/publications/2026/consumers-trust-ai-to-buy-better-brands-must-adapt` - B2B buyers trust AI less than marketers think - MarTech (Gartner survey coverage), May 2026 - `martech.org/b2b-buyers-trust-ai-less-than-marketers-think/` - Introducing Search Generative AI performance reports in Search Console - Google Search Central Blog, June 2026 - `developers.google.com/search/blog/2026/06/gen-ai-performance-reports` - LLaMAs Have Feelings Too: Unveiling Sentiment and Emotion Representations in LLaMA Models Through Probing - arXiv, May 2025 - `arxiv.org/abs/2505.16491` - Tracking Brand Mentions in AI Chatbots (Feb 2026 data) - rocketblue - `rocketblue.ai/articles/tracking-brand-mentions-in-ai-chatbots-a-comprehensive-guide-to-monitoring-brand-presence-in-chatgpt-responses-feb-2026-data/` --- ## The 10 Best Open-Source LLMs in August 2026 (Ranked) > The best open source LLMs of August 2026, ranked: GLM-5.2, DeepSeek V4, Kimi K3 and more, with licenses, real hardware needs, and every claim dated and verified. - Canonical: https://geotoolbox.ai/blog/best-open-source-llms - Published: 2026-07-22 · Updated: 2026-08-14 Picking the best open source LLM in 2026 has a specific problem: the answer changes monthly, and most of the rankings you will find were right when they were written and wrong by the time you read them. This one is dated on every claim, against the update stamp at the top of this page (the Kimi K3 ranking and license were last re-checked on August 1, 2026, and the Qwen3.8-Max status on August 8, 2026), and honest about which numbers come from vendor decks versus verified model cards. Ten models made the cut. Their licenses differ more than their benchmark scores do, and only a handful run on hardware you can actually own. ## The Best Open-Source LLMs in July 2026, at a Glance The short answer: **GLM-5.2** is the best open-source LLM you can download and run today, **DeepSeek V4** is the best value, and **Kimi K3** is the new benchmark leader of the open field, the top open-weight model on Artificial Analysis's Intelligence Index, with downloadable weights public since July 27, 2026. Its 2.8-trillion-parameter scale keeps it out of self-hosting reach for almost everyone, though, so expect the #1 slot to stay contested now that the weights have landed. Which one you should use depends on the job, the license terms you can live with, and the hardware you have. Here is the full field, verified against official model cards and launch documentation on July 22, 2026.
ModelMakerParameters (total / active)ContextLicenseBest for
GLM-5.2Z.ai753B / 40B1MMITBest overall, agentic coding
DeepSeek V4 Pro / FlashDeepSeek1.6T / 49B and 284B / 13B1MMITValue, world knowledge
Kimi K2.6 / K3Moonshot AI1T / 32B and 2.8T / 104B (K3)256K / 1MModified MIT (K2.6); custom Kimi K3 LicenseAgent swarms, frontier scale
MiniMax M3MiniMax428B / 23B1MMiniMax CommunityMultimodal agents
Qwen3.5 / 3.6 familyAlibaba0.8B to 397B / 17B262K+Apache 2.0License freedom, multilingual
MiMo-V2.5-ProXiaomi1.02T / 42B1MMITToken efficiency
Llama 4 Scout / MaverickMeta109B / 17B and 400B / 17Bup to 10M claimedLlama 4 CommunityWestern default, long context
Gemma 4GoogleE2B to 31B (4 sizes)256K (31B)Apache 2.0Local and edge hardware
gpt-oss-120bOpenAI117B / 5.1B128KApache 2.0Permissive Western reasoning
Nemotron 3NVIDIA550B / 55B (Ultra)1MOpenMDW-1.1Closest to truly open
Two things stand out in that table. First, six of the ten entries come from Chinese labs, and they hold most of the top benchmark slots. Second, "open" spans four meaningfully different license families, and the differences bite in production. One caveat before the rankings: AI engines themselves disagree on the #1 pick. When we ran this exact question through five engines in July 2026 (ChatGPT, Gemini, Perplexity, Claude, and ChatGPT's search mode), four named GLM-5.2 the best overall open model and one ranked DeepSeek V4 Pro first. The gap between them is small enough that your use case, not the leaderboard, should break the tie. ## Open Source vs Open Weight, in One Minute Almost every model on this list is **open weight**, not open source. The distinction matters more than the marketing suggests. Open weight means you can download the model parameters, run them on your own hardware, and usually use them commercially. It does not mean you can see the training data or reproduce the model. The [Open Source Initiative's Open Source AI Definition](https://opensource.org/ai/open-source-ai-definition) requires three things: detailed data information, the complete training code, and the parameters, all under open terms. Weights alone do not clear that bar. By the strict OSI reading, GLM-5.2, DeepSeek V4, and Llama 4 are all open weight. The models that genuinely qualify as open source, like AI2's OLMo or EleutherAI's Pythia, publish training data and recipes but do not top capability leaderboards. NVIDIA's Nemotron 3 comes closest to bridging the two: it ships weights alongside training datasets and recipes, which is why it earns a slot in these rankings despite mid-pack scores. For the full breakdown of what each label legally means, see our guide to [open weights vs open source](https://geotoolbox.ai/blog/open-weights-vs-open-source). The practical takeaway for this article: every model ranked below has downloadable open weights today, K3's included since July 27, and we flag the license catch on each one. ## How We Ranked These (and How to Read LLM Benchmarks) Every ranking below rests on official model cards and launch documentation, cross-checked against independent leaderboards that publish their setups, with every claim re-verified the week of July 22, 2026. That last part matters more than it should, and here is why. **Benchmark numbers are easy to game and easier to misread.** Three traps account for most of the bad comparisons you will see in open-source LLM lists: **Trap 1: Verified vs Pro.** SWE-bench Verified and SWE-bench Pro are different tests. GLM-5.2 scores 62.1% on [SWE-bench Pro per its official model card](https://huggingface.co/zai-org/GLM-5.2), while several rivals advertise 80%-class scores on the easier Verified set. Put those in one column and GLM-5.2 looks like it lost. It did not; the tests changed. **Trap 2: single-attempt vs multi-attempt.** Kimi K2 famously reported 65.8% on SWE-bench Verified in one attempt and 71.6% when allowed retries. Both are real numbers. Only one of them is comparable to a single-attempt score from another lab. **Trap 3: tools on vs tools off.** Humanity's Last Exam scores swing by 10 or more points depending on whether the model can search and run code. GLM-5.2's own card reports HLE at 40.5; leaderboards that allow tools report the same model in the mid-50s. Neither number is wrong, but putting them in one column is. There is also the blunter problem the r/LocalLLaMA crowd calls benchmaxxing: labs tune models toward the public tests. The pattern shows up as a model that tops SWE-bench yet fumbles your actual repo. Treat every score as a screening filter, then run your own task suite before committing. A model can ace a benchmark format and still fail a slightly reworded version of the same problem. With those caveats stated, the rankings. ## The Rankings: 10 Best Open-Weight Models Right Now ### 1. GLM-5.2 (Z.ai): Best Overall **753B total / 40B active MoE, 1M context, MIT license, announced June 13, 2026 with weights on Hugging Face three days later.** GLM-5.2 is the consensus pick for best open model of mid-2026, and the [official model card](https://huggingface.co/zai-org/GLM-5.2) backs it with a 91.2% GPQA Diamond score and the strongest terminal and long-horizon coding results in the open field. The MIT license is as clean as licensing gets: commercial use, modification, redistribution, no thresholds. The 1M-token context window is the practical centerpiece. Coding agents can hold a mid-sized repository in memory without constant compaction, and Z.ai reports its IndexShare optimization cuts long-context compute by roughly 2.9x at the full window. **Watch for:** no image input, and self-hosting is a serious project. Even aggressively quantized community builds still weigh hundreds of gigabytes. Most teams will run it through an API provider, not a garage rig. We cover the full specs in [what GLM-5.2 is](https://geotoolbox.ai/blog/what-is-glm-5-2). ### 2. DeepSeek V4 Pro and V4 Flash: Best Value **1.6T / 49B (Pro) and 284B / 13B (Flash), 1M context, MIT license.** DeepSeek's [V4 pair](https://geotoolbox.ai/blog/deepseek-v4) covers both ends of the budget. Pro is the flagship for reasoning and coding; Flash gets surprisingly close when given a larger thinking budget, at a fraction of the serving cost. Both were pre-trained on over 32T tokens and use a compressed-attention design that cuts KV-cache pressure dramatically at long context. V4 Pro's quietest advantage is factual knowledge: its model card publishes SimpleQA-Verified results that lead the open-model comparisons it reports. Strong reasoners that also know things are rarer than the leaderboards suggest. The company that triggered the original "DeepSeek moment" is still the value benchmark, and [DeepSeek pricing](https://geotoolbox.ai/blog/deepseek-pricing) remains the number every rival gets measured against. Background on the lab and the R1 story is in our [DeepSeek explainer](https://geotoolbox.ai/blog/what-is-deepseek). **Watch for:** Pro API throughput has been inconsistent since launch while DeepSeek scales serving capacity, so latency can spike. ### 3. Kimi K2.6 and K3 (Moonshot AI): Best for Agents, New Frontier Flagship **K2.6: ~1T / 32B, 256K context. K3: 2.8T parameters, 1M context, launched July 16, 2026.** K2.6 is the proven workhorse: its model card documents agent swarms of up to 300 sub-agents across 4,000 coordinated steps, built for long autonomous sessions. K3 is the headline: [the largest open-weight model ever released](https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems), a 2.8T-total / 104B-active MoE roughly 75% bigger than DeepSeek V4 Pro. It is now the top open-weight model on an independent leaderboard, scoring **57 on [Artificial Analysis's Intelligence Index](https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5)** — comparable to Claude Opus 4.8 and GPT-5.5, and ahead of every other open model (GLM-5.2 at 51, DeepSeek V4 Pro at 44). Its launch benchmarks trade blows with the top proprietary systems, including a vendor-reported, field-leading 91.2 on BrowseComp. One date matters here: K3's weights went public on **July 27, 2026**. Before that it was API-and-app only, still priced at [$3 per million input tokens and $15 per million output](https://venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems). Our [Kimi K3 breakdown](https://geotoolbox.ai/blog/what-is-kimi-k3) tracks what is verified versus vendor-claimed, and our [how to run Kimi K3 locally](https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally) guide covers whether you can actually self-host a 2.8T model (mostly, no). **Watch for:** the license. K2.6 ships under Moonshot's [Modified MIT](https://huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE): products with more than 100 million monthly active users or more than $20 million in monthly revenue must prominently display "Kimi K2.6" in their UI. Trivial for most teams, a real conversation for consumer apps at scale. K3 shipped under its own [custom Kimi K3 License](https://huggingface.co/moonshotai/Kimi-K3) rather than the Modified MIT some expected: commercial use is permitted, products past 100 million MAU or $20 million in monthly revenue must display "Kimi K3" in the UI, and any company running a model-as-a-service business on K3 must sign a separate agreement with Moonshot once its total revenue, counting affiliates, passes $20 million over any 12 months. Runtime support was faster than usual for a new attention design: vLLM had day-zero support on July 27 and GGUF and Ollama quantizations were live at the weight drop, not weeks later. ### 4. MiniMax M3: Best Multimodal Agent Model **428B / 23B MoE, 1M context via MiniMax Sparse Attention.** M3 was trained mixed-modality from the start, so it takes image and video input natively, which most of its open rivals still cannot. MiniMax pitches it on endurance: internal demos show day-scale autonomous runs on tasks where rival models stall early, and the MSA design exists precisely to make those very long sessions affordable. **Watch for:** the MiniMax Community License carries commercial-use conditions, and the model's reasoning scores rest partly on aggregator leaderboards that weight tests differently than the individual benchmarks do. Read the license and the fine print on both. ### 5. Qwen3.5 and Qwen3.6 (Alibaba): Best License Freedom and Language Coverage **0.8B to 397B / 17B active, 262K-token context window extendable toward 1M, Apache 2.0.** Qwen is the ecosystem play. One family covers phone-sized 0.8B models up to the 397B Qwen3.5 flagship, all under Apache 2.0, the most permissive license in common use: no attribution thresholds, no user caps, nothing to renegotiate when your product grows. The flagship reasons across text, images, video, and documents in one framework and covers 200+ languages, which no other open family matches. The generations also turn over fast: the April 2026 **Qwen3.6** wave added a 27B dense model and a 35B-total, 3B-active MoE that punch far above their size, and they, not the Qwen3.5 flagship, are the family's current picks for consumer hardware. Full lineup in our [Qwen guide](https://geotoolbox.ai/blog/what-is-qwen). **Watch for:** running the flagship at full context demands a serious multi-GPU memory budget. The family's breadth is the point; buy the size you can serve. And a naming note: Alibaba's 2.4T-parameter **[Qwen3.8-Max](https://geotoolbox.ai/blog/qwen3-8-max)** shipped on August 3, 2026, but as the closed Max API line ($2 / $6 per million tokens), not an open release; its open weights landed on August 12, 2026, but as a text-only variant under a bespoke "Qwen3.8-Max" license rather than a permissive one, so it sits outside the permissively-licensed line this ranking covers. Within that line, Qwen3.5 remains the open flagship. ### 6. MiMo-V2.5-Pro (Xiaomi): The Token-Efficiency Dark Horse **1.02T / 42B MoE, MIT license, trained on 27T tokens.** Xiaomi is the least-discussed lab on this list and arguably the most efficient. MiMo-V2.5-Pro matches frontier open rivals on coding-agent tasks while spending markedly fewer tokens per run in Xiaomi's own evaluations, which compounds into real money across long agentic workloads. Its hybrid attention design holds performance past 512K tokens where its predecessor collapsed outright. **Watch for:** ecosystem maturity. Tooling, quantized community builds, and provider support all trail the bigger names, so expect more integration work. ### 7. Llama 4 Scout and Maverick (Meta): The Western Default **109B / 17B and 400B / 17B MoE, multimodal, context windows advertised up to 10M tokens.** Llama remains the infrastructure layer of the open ecosystem: the widest tooling support, the most fine-tunes, the most deployment guides. Scout is the efficiency pick for local development and edge work; Maverick is the production generalist with strong multilingual scores. Llama also powered [Meta AI](https://geotoolbox.ai/blog/what-is-meta-ai), Meta's consumer assistant, until the company moved that app to its proprietary Muse Spark model in 2026. **Watch for:** two things. The Llama 4 Community License is not OSI-approved and cuts off free commercial use at 700M monthly active users. And treat the 10M-token context as a marketing ceiling, not a working spec; long-context quality falls off well before advertised limits across the industry, so size your workloads to a far smaller usable window. ### 8. Gemma 4 (Google): Best for Local and Edge Hardware **Four sizes from Effective-2B to 31B dense, 256K context on the 31B, Apache 2.0.** Gemma 4 is the answer to "what can I actually run?" The family ladders cleanly: phone-class E2B and E4B variants that uniquely accept audio input, a 26B-A4B mixture of experts that activates only a 4B-class slice per token, and a 31B dense flagship whose unquantized bf16 weights fit a single 80GB H100, with quantized builds running on consumer GPUs. Google also fixed the licensing complaint: where Gemma 3 shipped under custom terms, Gemma 4 is Apache 2.0. **Watch for:** a 31B dense model does not compete with the trillion-parameter MoE tier on hard reasoning, and the 31B variant drops the audio support the small ones have. This is the local tier, not the frontier tier. ### 9. gpt-oss-120b (OpenAI): The Permissive Western Reasoner **117B / 5.1B active, 128K context, Apache 2.0.** OpenAI's open-weight release is easy to overlook next to the Chinese giants, and shouldn't be. With only 5.1B active parameters it is one of the cheapest strong reasoners to serve, it was built for reasoning and tool use, and the Apache 2.0 license plus a US-origin lab matters to procurement teams that cannot ship a Chinese model, fair or not. (The safety side of that question has [its own article](https://geotoolbox.ai/blog/chinese-ai-models-compared).) **Watch for:** it trails the frontier open models badly on broad-knowledge tests, and 128K context is now the smallest window on this list. ### 10. Nemotron 3 (NVIDIA): The Closest Thing to Truly Open **550B / 55B hybrid Transformer-Mamba (Ultra), 1M context, OpenMDW-1.1 license.** Nemotron earns its slot on transparency rather than peak scores. NVIDIA publishes weights, training datasets, and recipes, the closest any frontier-adjacent lab comes to the OSI's actual open-source bar. For teams that need real visibility into training data, or want to build on a model whose data provenance is documented, it is effectively the only option at this scale. **Watch for:** benchmark performance sits mid-pack, and the hybrid architecture has thinner community tooling than the standard MoE stacks. **Near misses.** Three absences are deliberate. Mistral's open line today is the Devstral 2 coding specialists plus [Mistral Large 3](https://huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512), a 675B Apache 2.0 generalist that predates this year's frontier wave. Microsoft's Phi-4 family owns the tiny, no-GPU end rather than the flagship fight. And OLMo, covered above, wins on openness rather than capability. All three are worth knowing; none displaces an entry above at its own job. ## Which Open Model for Which Job
You needUseWhy
Best overall / agentic codingGLM-5.2Top open scores on long-horizon coding, MIT, 1M context
Best value per dollarDeepSeek V4 FlashNear-Pro reasoning at a fraction of the price, MIT
Agent swarms and long autonomous runsKimi K2.6300 sub-agent orchestration across 4,000 coordinated steps
Multimodal agents (image, video, computer use)MiniMax M3Native mixed-modality training
Cleanest license at any sizeQwen3.5 or Gemma 4Apache 2.0, no thresholds
A single 24GB consumer GPUGemma 4 26B-A4B or a mid-size Qwen3.6Built for that hardware class
Frontier scale, newest weightsKimi K3 (from July 27)Top open-weight model on the AA Intelligence Index (57); 2.8T parameters
Hardest reasoning problemsDeepSeek V4 Pro (Think Max) or GLM-5.2Top open reasoning scores; adaptive thinking modes
Auditable training dataNemotron 3Data and recipes published, not just weights
And the free-tier answer, since it is one of the most-asked questions: every model above is free to download once weights are public. What costs money is the compute. If "free" means "free to chat with," the hosted apps for Kimi, DeepSeek, and Qwen all have no-cost tiers. ## Can You Actually Run These? The Self-Host Reality Check Here's the part most rankings skip: for the frontier tier of this list, "downloadable" and "runnable" are different claims. **The VRAM arithmetic.** A model needs roughly 2GB of memory per billion parameters at FP16, and roughly 0.5GB per billion at 4-bit quantization, plus KV-cache overhead that grows with context length and concurrency (modest for local chat, huge at long contexts). An 8GB GPU tops out around a 7B-8B model at 4-bit. A 24GB card handles the 27B-32B dense class. Nothing consumer-grade touches the trillion-parameter tier. **The MoE trap.** This is the single most misunderstood fact in self-hosting. A mixture of experts model like Kimi K2.6 activates only 32B parameters per token, but **all 1T parameters must sit in memory**, because any token can route to any expert. Active parameters set your speed and per-token cost. Total parameters set your memory bill. Community builds of K2.6-class models need hundreds of gigabytes of combined RAM and VRAM even at extreme quantization. **Quantization is the lever, with limits.** Q4_K_M, the standard 4-bit GGUF format, cuts memory roughly 70-75% with only modest quality loss on most dense models. But MoE models degrade less gracefully under aggressive quantization, and reasoning quality tends to fall before chat quality does. If a model's edge is careful multi-step reasoning, test the quant before trusting it. **The stack.** For prototyping, [Ollama](https://ollama.com/) is the fastest path: one command pulls a pre-quantized build and serves it behind an OpenAI-compatible endpoint. For production concurrency, [vLLM](https://docs.vllm.ai/en/latest/) is the standard, with PagedAttention and continuous batching to keep GPUs saturated; SGLang is its main rival. Whichever you pick, judge the setup on the two numbers users actually feel, time to first token and inter-token latency, not peak throughput. For a step-by-step walkthrough from install to your first chat, see our guide on [how to run an LLM locally](https://geotoolbox.ai/blog/run-llm-locally).
HardwareWhat it runsExamples
Laptop / 8GB GPU2B-8B at 4-bitGemma 4 E2B-E4B, small Qwen3.5/3.6
24GB consumer GPU12B-32BGemma 4 26B-A4B (quantized), Qwen3.6 mid-tier
128GB Mac Studio / workstation~100-300B MoE at 3-bit-class quantsDeepSeek V4 Flash class, gpt-oss-120b
Multi-GPU server (300GB+)The frontier tier, quantizedGLM-5.2, Kimi K2.6, DeepSeek V4 Pro
API onlyEverything, no capexAll of the above via provider or official APIs
**So what does running one really cost?** For the frontier-tier models, usually more than the API. A single used H100-class GPU runs five figures before power, and the frontier tier needs several. Meanwhile per-token API prices for open models are aggressive: DeepSeek V4 Flash sells for cents per million tokens, and even Kimi K3 launched at $3 per million input. For scale: a workload generating a full million output tokens every day on K3, the priciest model here, runs about $450 a month at its $15-per-million output rate. Self-hosting wins on data privacy (a self-hosted model sends no prompts to anyone's API, which removes the data-residency half of the Chinese-model worry, though license terms and security review still apply), offline operation, and fine-tuning freedom, and at sustained very high volume. It rarely wins on cost alone; our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) breakdown shows how quickly hosted rates undercut owned hardware for typical workloads. Run the math on your actual token volume before buying anything. ## Why Every "Best Open LLM" List Disagrees Search this topic and you will find current, well-ranked articles confidently recommending three different generations of the same model line as "the latest." One top-ranking guide's newest entry is a model two generations old. That is velocity. The open-weight field now ships a new flagship roughly every four to six weeks, and static listicles rot in months. The lineage table below is the fix. It names the current flagship per lab, verified July 22, 2026, next to the superseded versions you will still see recommended elsewhere.
LabCurrent flagship (July 22, 2026)Superseded, still widely recommended
Z.ai (Zhipu)GLM-5.2 (June 2026)GLM-5.1, GLM-5
Moonshot AIKimi K3 (July 16; weights July 27) + K2.7-Code (June 2026, coding fork)K2.6, K2.5, K2 Thinking, K2
DeepSeekV4 Pro / V4 Flash (April 2026)V3.2, V3-0324, R1
Alibaba (Qwen)Qwen3.6 series (Apr 2026); Qwen3.5-397B stays the flagship by size*Qwen3-Coder-480B, Qwen3 VL 235B, Qwen3
MiniMaxM3 (June 2026)M2.5, M2
MetaLlama 4 (April 2025)Llama 3.3, 3.1
GoogleGemma 4 (April 2026)Gemma 3
XiaomiMiMo-V2.5-ProMiMo-V2-Pro
*Qwen's naming runs two tracks (open-weight releases and the closed Max API line), which is itself a recurring source of listicle confusion. Case in point: the 2.4T Qwen3.8-Max shipped August 3, 2026 as the closed Max API line, with open weights delivered August 12, 2026 as a text-only variant under a bespoke license, a partial opening that still sits outside the permissively-licensed line this ranking covers. Two data points explain the churn. [Epoch AI measured](https://epoch.ai/data-insights/open-weights-vs-closed-weights-models) frontier open-weight models trailing the best closed models by an average of about three months as of late 2025; [its May 2026 update](https://epoch.ai/data-insights/open-closed-eci-gap) puts the lag at roughly four months. Separately, [Stanford's AI Index](https://hai.stanford.edu/ai-index/2025-ai-index-report) found the open-versus-closed performance gap on some benchmarks collapsing from 8 to 1.7 percentage points in a single year, and the spread between the globally top-ranked ten models narrowing from 11.9 to 5.4 points. A field that turns over in months makes every month-old ranking partially wrong, including, eventually, this one. That is why every claim here carries a date, the same rule geotoolbox applies to its own engine scans.
![Timeline of flagship open-weight LLM releases from January 2025 to July 2026 across DeepSeek, Moonshot AI, Z.ai, MiniMax, Alibaba Qwen, Meta and Google, with Kimi K3 highlighted in July 2026.](/blog/best-open-source-llms/open-weight-release-race-2025-2026.png)
Six of these labs shipped a new flagship in the last six months. Dates from official model cards and launch coverage, verified July 22, 2026.
## What the Open-Weight Wave Means for Your AI Visibility If you work in marketing or SEO, this model race concerns you directly. Every model on this list is also an answer engine that describes brands, recommends products, and cites sources. Open weights accelerate that in a specific way: they get embedded into products silently. When a coding tool, a CRM, or a search feature quietly ships GLM-5.2 or a Qwen fine-tune under the hood, it inherits that model's picture of your brand, and you will never see a referrer string telling you so. The practical response is the same discipline this article applies to models: verify instead of assume. Check what the engines actually say about your brand, measure which of your pages they cite, and re-check when a new model generation lands, because a version bump swaps out the training data, the retrieval behavior, and often the citation habits in one move. A brand answer that held for Kimi K2.5 is not guaranteed to survive K3. That monitoring is what geotoolbox is built for. If you want to know whether your site is even readable by the crawlers feeding these models, start with the free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness); it takes a minute and flags the blockers that keep your brand out of AI answers regardless of which model wins the month. ## Frequently Asked Questions ### What is the strongest open-source LLM right now? It depends on what you weigh. On raw benchmarks, Kimi K3 is now the top open-weight model: as of August 1, 2026 it scores 57 on Artificial Analysis's Intelligence Index, ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44), and its weights went public on July 27. But K3's 2.8-trillion-parameter size makes it effectively impossible to self-host, so GLM-5.2, with a 91.2% GPQA Diamond score, a clean MIT license, and a memory footprint you can actually fit on a multi-GPU server, remains the strongest open model most teams can download and run today. ### Is there a truly open-source LLM? Yes, but not at the frontier. The OSI bar requires three things published: data information, training code, and parameters. AI2's OLMo and EleutherAI's Pythia clear it; NVIDIA's Nemotron 3 comes closest among large models. The benchmark leaders on this page, DeepSeek included, are open weight, not OSI open source. ### Is there an open-source LLM better than ChatGPT? On specific coding and agentic benchmarks, yes: GLM-5.2, DeepSeek V4 Pro, and Kimi K3 (weights July 27) match or beat particular proprietary model versions on particular tests. As a product, ChatGPT bundles tools, search, and model routing that no single open model replicates, and the strongest closed models still hold an overall edge. Epoch AI put the average capability lag at about four months in its May 2026 update. ### How much does it cost to run an open-source LLM? The weights are free; the compute is not. A 7B-12B model runs on a consumer GPU you may already own. The 24GB-GPU class covers models up to roughly 32B. Frontier models like GLM-5.2 or Kimi K2.6 need hundreds of gigabytes of memory, which means a multi-GPU server or a hosted API. For most workloads, per-token API pricing (from cents per million tokens for DeepSeek V4 Flash) is cheaper than owning hardware. ### What is the best open-source LLM for coding? GLM-5.2 for long-horizon agentic coding, on the strength of its terminal and repository-scale benchmark leads. DeepSeek V4 is the strongest value pick, and Qwen's coder variants offer the best efficiency per active parameter under Apache 2.0. Moonshot's K2.7-Code fork and retry-friendly agent frameworks close much of any remaining gap. ### Can I download Kimi K3 yet? Moonshot AI released K3's public weights on July 27, 2026; before that it was available through the Kimi app and API only. Its 2.8T-parameter size means self-hosting will demand multi-node hardware even at aggressive quantization; expect the practical route to remain hosted access. ## Sources - China's Moonshot AI releases Kimi K3, the largest open-source model ever - VentureBeat, July 16, 2026 - `venturebeat.com/technology/chinas-moonshot-ai-releases-kimi-k3-the-largest-open-source-model-ever-rivaling-top-u-s-systems` - Kimi K3 achieves #3 in the Artificial Analysis Intelligence Index, comparable to Opus 4.8 and GPT-5.5 - Artificial Analysis, July 17, 2026 - `artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5` - Kimi K3 model card and license - Moonshot AI on Hugging Face - `huggingface.co/moonshotai/Kimi-K3` - The Open Source AI Definition v1.0 - Open Source Initiative - `opensource.org/ai/open-source-ai-definition` - How far behind are open models? - Epoch AI, October 30, 2025 - `epoch.ai/data-insights/open-weights-vs-closed-weights-models` - The 2025 AI Index Report - Stanford HAI - `hai.stanford.edu/ai-index/2025-ai-index-report` - GLM-5.2 official model card - Z.ai on Hugging Face - `huggingface.co/zai-org/GLM-5.2` - Kimi K2.6 Modified MIT license text - Moonshot AI on Hugging Face - `huggingface.co/moonshotai/Kimi-K2.6/blob/main/LICENSE` - Mistral Large 3 675B Instruct model card - Mistral AI on Hugging Face - `huggingface.co/mistralai/Mistral-Large-3-675B-Instruct-2512` - Open models lag update (ECI gap) - Epoch AI, May 2026 - `epoch.ai/data-insights/open-closed-eci-gap` - vLLM documentation - `docs.vllm.ai/en/latest/` - Ollama - `ollama.com` --- ## Gemini 3.6 Flash vs 3.5 Flash-Lite: Which to Use (2026) > Google shipped Gemini 3.6 Flash and 3.5 Flash-Lite on July 21, 2026. Here's what each costs, how they benchmark, and which cheap Gemini model to actually use. - Canonical: https://geotoolbox.ai/blog/gemini-3-6-flash-vs-3-5-flash-lite - Published: 2026-07-22 · Updated: 2026-08-08 On July 21, 2026, Google made two new Gemini Flash models generally available: **Gemini 3.6 Flash** and **Gemini 3.5 Flash-Lite**. Both run in the Gemini API today, both carry a 1-million-token context window, and both are aimed at the cheap, high-volume end of the lineup. The specs are the easy part. The harder question is the one developers are actually asking: with several overlapping low-cost Gemini models now live, which one do you use, and did the newest Flash quietly get more expensive? Short version: Flash-Lite for high-volume work, 3.6 Flash for agentic coding, and a flagship when the code gets hard. The reasons are below. ## What Google Actually Shipped on July 21 Google announced three models, not two. The third is **Gemini 3.5 Flash Cyber**, a security-tuned system for finding and patching vulnerabilities inside Google's CodeMender, limited to governments and trusted partners, so most developers can ignore it for now. The two you can actually call are 3.6 Flash and 3.5 Flash-Lite, live in the Gemini API, AI Studio, and Android Studio, with rollouts into the Gemini app and Google Search. The naming trips people up, and for good reason. Flash jumped to **3.6** while Flash-Lite stayed on the **3.5** generation. There is no 3.6 Flash-Lite. Google framed the release as expanding its Gemini 3.5 line, with Flash getting a point-bump to 3.6 as the direct successor to the 3.5 Flash it shipped in May. The flagship most people were waiting for, Gemini 3.5 Pro, is still testing with partners (the Pro model you can use today is still labeled 3.1 Pro Preview), and DeepMind has confirmed it already started pre-training [Gemini 4](https://9to5google.com/2026/07/21/gemini-3-6-flash-launch/). In other words, this is a mid-cycle refresh of the cheap tier, not the Pro release. If you are new to the lineup, our [Google Gemini overview](https://geotoolbox.ai/blog/what-is-gemini) maps how the models fit together. ## Gemini 3.6 Flash: What Changed Gemini 3.6 Flash (`gemini-3.6-flash`) is the new workhorse. It replaces 3.5 Flash and is priced at **$1.50 per million input tokens and $7.50 per million output tokens**, down from 3.5 Flash's $9.00 output. It keeps the 1M-token context window with a 64k output cap, takes text, image, video, audio, and PDF input, and ships with thinking controls and built-in tools including Computer Use. Two changes matter more than the rest. First, token efficiency: Google says 3.6 Flash uses **17% fewer output tokens** than 3.5 Flash on the Artificial Analysis Index, and on specific agentic workloads the gap is larger. On DeepSWE, a coding benchmark, it used roughly 97,000 output tokens per task versus 276,000 for its predecessor. Fewer output tokens means a lower bill even at the same rate, because output is where the cost concentrates. Second, and easy to overlook, the knowledge cutoff moved from **January 2025 to March 2026**. That 14-month jump is arguably the most practical upgrade in the release. A model that knows about recent framework releases, pricing changes, and product launches gives fewer confidently wrong answers about anything from the last year. On its own benchmarks, 3.6 Flash beats 3.5 Flash across the board: DeepSWE coding rises to 49% from 37%, GDPval-AA v2 knowledge work to 1421 from 1349, and OSWorld-Verified computer use to 83.0% from 78.4%, per [Google's launch post](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/). Google's model page adds SWE-Bench Pro at 58.7% versus 55.1%. These are vendor-reported numbers. The independent picture is more mixed, and we get to it below. ## Gemini 3.5 Flash-Lite: The New Budget Tier Gemini 3.5 Flash-Lite (`gemini-3.5-flash-lite`) is the cheap, fast option: **$0.30 per million input tokens and $2.50 per million output tokens**, the same rate as 2.5 Flash but newer. It generates around **350 output tokens per second**, which Google calls the fastest in the Gemini 3.5 series, and it defaults to minimal thinking so it stays quick on high-volume traffic. The surprise is how capable it is for the price. On Google's benchmarks, Flash-Lite outperforms the older Gemini 3 Flash on SWE-Bench Pro (54.2% versus 49.6%) and on OSWorld-Verified (74.0% versus 65.1%), and it posts large gains over the previous 3.1 Flash-Lite: Terminal-Bench 2.1 jumped from 31% to 54%, and long-context retrieval on GDM-MRCR v2 rose from 60.1% to 72.2%. For a model in the cheapest production tier, beating a full Flash model on agentic benchmarks is not what the name suggests. Flash-Lite is built for translation, classification, document extraction, routing, agentic search, and fast user-facing features, the work where you send millions of tokens through and care most about speed and cost. ## Gemini Flash Pricing, Compared Set against the rest of the low-cost lineup, at Google's official Standard-tier rates per million tokens:
ModelInput / 1MOutput / 1MPositioning
Gemini 3.6 Flash$1.50$7.50New workhorse: agentic coding, multimodal
Gemini 3.5 Flash-Lite$0.30$2.50New budget tier: high-volume, low-latency
Gemini 3.5 Flash$1.50$9.00Superseded by 3.6 Flash
Gemini 2.5 Flash$0.30$2.50Prior general-purpose cheap model
Gemini 2.5 Flash-Lite$0.10$0.40Still the cheapest Gemini tier
Gemini 3.1 Pro Preview$2.00$12.00Pro tier (doubles above 200k context)
A common complaint after launch was that "Gemini Flash now costs 25 times more than Gemini 1.5 Flash." That is roughly true for the Flash tier: 3.6 Flash's $7.50 output is far above the sub-dollar output rate of the old 1.5 Flash. But it misreads what happened. "Flash" moved upmarket toward being a Pro replacement, and the cheap end shifted down a rung to **Flash-Lite**. If you want rock-bottom cost, 3.5 Flash-Lite at $2.50 output, or 2.5 Flash-Lite at $0.40, is the tier you want, not the model still labeled Flash. The gap between those two Lite tiers is real money. Running one million short classification calls at roughly 500 output tokens each costs about $1,250 in output on 3.5 Flash-Lite, against about $200 on 2.5 Flash-Lite. The older, cheaper 2.5 Flash-Lite is the right pick when the task is simple enough that quality headroom does not matter; step up to 3.5 Flash-Lite when it starts making mistakes you have to clean up. Two more levers cut either bill further on high-volume work: batch mode, which trades real-time responses for a lower asynchronous rate, and context caching, which charges less to reuse a shared prompt prefix. One trap worth flagging: some resellers list their own rates. A third-party platform quoting Flash-Lite at $0.25 input and $1.50 output is not quoting Google, whose official rate is $0.30 and $2.50. Price against the [official Gemini API rates](https://ai.google.dev/gemini-api/docs/pricing), and check our [Gemini API pricing breakdown](https://geotoolbox.ai/blog/gemini-api-pricing) for the free tier and the meters that inflate a real bill. The token-efficiency angle changes the real math on 3.6 Flash. Its output rate already dropped to $7.50 from $9.00, and it emits about 17% fewer output tokens on top of that, so the cost of finishing a given task falls on two fronts, not just the sticker rate. On output-heavy, high-volume work, that is where the savings actually show up. ## The Honest Capability Read Google's benchmark table shows gains everywhere. Independent testing tells a flatter story, and it is worth holding both in view. On the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gemini-3-6-flash), a composite score, Gemini 3.6 Flash lands at **50**, the same score as the 3.5 Flash it replaced. Flash-Lite scores 36. Read together with the per-benchmark wins, the takeaway is specific: 3.6 Flash is faster, cheaper, and more token-efficient than 3.5 Flash, but not measurably smarter on general reasoning. The gains landed in cost and throughput, not in the aggregate intelligence score. Against the frontier, it trails. [DataCamp's benchmark roundup](https://www.datacamp.com/blog/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber) has frontier models ahead on hard coding: Grok scores 64.7% on SWE-Bench Pro to 3.6 Flash's 58.7%, and Claude Sonnet 5 leads on machine-learning engineering (66.9% versus 63.9%). Those are each lab's reported figures rather than one shared harness, so treat them as directional. The same roundup notes 3.6 Flash wins on long-context retrieval and chart reasoning but trails on pure coding, which is why the launch-day social reaction was harsh. That is the right lens, though: Flash is a mid-tier model, not a flagship. Judged against the frontier it looks middling; judged on price-per-task for high-volume agentic work, it is competitive. One practical wrinkle: latency. 3.6 Flash's throughput is strong, among the faster models Artificial Analysis measured at launch, but its time to first token ran well over 11 seconds in that testing, far above the roughly 3-second median for its price bracket. It can feel slow in interactive chat even while it streams bulk output quickly, so it fits batch and agent pipelines better than a live chatbot. If you are weighing it against Anthropic's or OpenAI's models, our [Claude vs Gemini comparison](https://geotoolbox.ai/blog/claude-vs-gemini) covers the trade-offs. ## Which Gemini Flash Model Should You Use? Match the model to the workload, not to the version number.
![Decision matrix for the Gemini Flash models. Use Gemini 3.6 Flash ($1.50 / $7.50 per million tokens) for agentic coding, multimodal, and multi-step reasoning. Use 3.5 Flash-Lite ($0.30 / $2.50) for high-volume extraction, classification, and routing, and for fast user-facing features where latency shows. Use 2.5 Flash-Lite ($0.10 / $0.40) for rock-bottom cost on simple tasks. Use a flagship model for frontier-grade code quality, since Flash trails Grok 4.5 and Sonnet 5 on hard coding.](/blog/gemini-3-6-flash-vs-3-5-flash-lite/gemini-flash-decision-matrix.png)
Which Gemini Flash model to use, by workload. Prices are Google's official per-million-token rates (input / output).
A pattern worth stealing for agent builds: use **3.6 Flash as a coordinating agent** that plans work and reviews results, and hand the parallel grunt work (file search, extraction, test writing) to multiple copies of **3.5 Flash-Lite**. One [hands-on review](https://acceleratedlogicai.com/blog/gemini-3-6-flash-and-3-5-flash-lite-review) that tested both found them dependable in repeated coding loops, though not something to trust blindly on critical production code. Keep a human on the path that matters. ## Migrating Your Gemini API Code Switching to the new models is mostly a model-ID change, but Gemini 3.x brings breaking changes that will bite if you copy an old config across. The mechanical steps: 1. Update the model string to `gemini-3.6-flash` or `gemini-3.5-flash-lite`. 2. Remove `temperature`, `top_p`, and `top_k`. These sampling parameters are [deprecated and ignored](https://dev.to/googleai/gemini-36-flash-35-flash-lite-developer-guide-i17) on Gemini 3.x, and future versions will reject them with an HTTP 400. 3. Replace `thinking_budget` with the `thinking_level` string enum (`minimal`, `medium`, or `high`). Set Flash-Lite to `minimal` for high-volume extraction, higher for tool-calling subagents. 4. Drop `candidate_count`, which is unsupported, and stop sending prefilled model turns, which now return a 400. Do not swap prod on faith. Run both models against 20 to 50 representative tasks and compare accuracy, latency, output length, and total cost before you commit. Watch specifically for the weaker frontend and UI generation some early testers flagged, and re-tune prompts where you see it. There is no forced deadline: Google has not announced a retirement date for 3.5 Flash or the 2.5 models, so migrate when the newer pricing and quality earn it, not because you have to. ## What the New Flash Models Mean for AI Search Visibility One detail matters more than the benchmark table if you care about being found in AI answers. Ask a web-connected engine about Gemini 3.6 Flash and it answers correctly. Ask a training-only model the same question and it denies the model exists. Queried a day after launch, Gemini 3.5 Flash itself replied, "Google has not announced or released models named Gemini 3.6 Flash," and suggested the user was confusing it with an older Gemini release. Anthropic's Claude said much the same. A Google model does not know about Google's newest model. That is not a Gemini quirk. It is how any model without live retrieval works: its training has a cutoff, and anything newer only reaches it through search or another connector. When someone asks an AI engine about your product, your pricing, or a feature you shipped last month, the answer rides whatever the model was trained on plus whatever it can pull in at that moment. If your current facts are not on a page a retrieval system can reach and lift, the model fills the gap with something older, or wrong. A common reason we see a brand's fresh information never reach an AI answer is a reachability gap on its own pages, not the model. If AI search is a channel you care about, the [free AI readiness check](https://geotoolbox.ai/tools/ai-readiness) shows whether AI crawlers can reach and parse your content, and our guide on [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) covers how to measure whether you are being cited. ## The Short Answer For high-volume, latency-sensitive work like extraction, classification, routing, and agentic search, use **3.5 Flash-Lite**: it is the cheaper of the new pair, the faster, and it beats the older Gemini 3 Flash on Google's reported benchmarks. For agentic coding, multimodal tasks, and multi-step reasoning, use **3.6 Flash**, where the token savings genuinely offset the higher rate. If you need the absolute floor on price for simple work, **2.5 Flash-Lite** is six times cheaper on output than 3.5 Flash-Lite. And if frontier code quality is the goal, Flash is the wrong tier, reach for a flagship. Pick by workload, retest on your own tasks, and price against Google's official rates, not a reseller's. ## Frequently Asked Questions ### Is Gemini Flash free? There is a free tier in Google AI Studio and the Gemini API with rate limits, and the consumer Gemini app has a free plan. Production API usage is paid per token. The free tier also comes with a data-use catch, which our [Gemini API pricing guide](https://geotoolbox.ai/blog/gemini-api-pricing) explains in full. ### What is the cheapest Gemini model? Gemini 2.5 Flash-Lite, at $0.10 per million input tokens and $0.40 per million output tokens, remains the cheapest Gemini model. Among the new July 2026 releases, 3.5 Flash-Lite is the cheapest at $0.30 input and $2.50 output. ### Is Gemini 3.6 Flash better than 3.5 Flash? On Google's benchmarks, yes, it wins on coding, knowledge work, and computer use. On the independent Artificial Analysis Intelligence Index it scores the same 50, so it is not measurably smarter overall. What it clearly is: cheaper output ($7.50 versus $9.00), 17% more token-efficient, and current to March 2026. ### Why does Gemini Flash cost more than older Flash models? The Flash tier moved upmarket. Gemini 3.6 Flash is positioned closer to a Pro replacement than the old 1.5 Flash was, so its $7.50 output looks steep next to sub-dollar legacy rates. The cheap end shifted to Flash-Lite, which is where high-volume, cost-sensitive work now belongs. ### What is the difference between Gemini Flash and Flash-Lite? Flash (3.6) is the higher-quality workhorse for coding, reasoning, and multimodal tasks at $1.50 / $7.50. Flash-Lite (3.5) is the cheaper, faster tier for high-volume extraction and low-latency features at $0.30 / $2.50, with lower intelligence but higher throughput. ### When will Gemini 3.5 Pro and Gemini 4 launch? As of the July 21, 2026 release, Gemini 3.5 Pro is still testing with partners and Google has only said broad availability is coming. DeepMind confirmed it has started pre-training Gemini 4 but gave no date. Our [Gemini 3.5 Pro tracker](https://geotoolbox.ai/blog/gemini-3-5-pro) follows what is confirmed. (Verified still accurate as of August 8, 2026, against Google's published model list.) ## Sources - Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber - Google, July 21 2026 - `blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber` - Gemini API Pricing - Google AI for Developers - `ai.google.dev/gemini-api/docs/pricing` - Gemini 3.5 Flash-Lite vs Gemini 3.6 Flash - Artificial Analysis - `artificialanalysis.ai/models/comparisons/gemini-3-5-flash-lite-vs-gemini-3-6-flash` - Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4 - 9to5Google, July 21 2026 - `9to5google.com/2026/07/21/gemini-3-6-flash-launch` - Gemini 3.6 Flash & 3.5 Flash-Lite Developer Guide - dev.to/googleai - `dev.to/googleai/gemini-36-flash-35-flash-lite-developer-guide-i17` - Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber - DataCamp - `datacamp.com/blog/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber` - Gemini Releases 3.6 Flash and 3.5 Flash-Lite: Are They Any Good? - Accelerated Logic AI - `acceleratedlogicai.com/blog/gemini-3-6-flash-and-3-5-flash-lite-review` --- ## How to Get Cited by AI: What Actually Works in 2026 > How to get cited by AI: what actually determines citations in ChatGPT, Perplexity, Gemini and AI Overviews - the evidence, the myths, and a 30-day plan. - Canonical: https://geotoolbox.ai/blog/how-to-get-cited-by-ai - Published: 2026-07-22 · Updated: 2026-07-22 Ask ChatGPT or Perplexity a question in your category and a handful of pages get named as sources. Everyone else is invisible. Figuring out how to get cited by AI is mostly a matter of separating what engines demonstrably reward from a growing pile of folklore. The short version: citations are chosen at the passage level, a reachability check almost everyone skips comes before any content work, every engine picks its sources differently, and measuring your progress takes more than one check. No tricks in here, just the parts with evidence behind them. ## What Counts as an AI Citation An **AI citation** is a linked reference to your page inside a generated answer: ChatGPT's source chips, Perplexity's numbered footnotes, the link cards beside a [Google AI Overview](https://geotoolbox.ai/blog/what-are-google-ai-overviews). It is not the same thing as a mention, and the difference decides what you optimize for. A [brand mention](https://geotoolbox.ai/glossary/brand-mention) names your brand in the answer text. An [AI citation](https://geotoolbox.ai/glossary/ai-citation) links your page as a source. A recommendation goes further and puts you forward as the answer. You can be mentioned without being cited (the engine learned about you from other people's pages) and cited without being mentioned (your data backs someone else's claim). Each combination tells you something different about your [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility). Why chase citations at all? The traffic is small, but it is not ordinary traffic. Semrush's [AI search traffic study](https://www.semrush.com/blog/ai-search-seo-traffic-study/) found the average AI search visitor is worth 4.4 times the average organic search visitor, measured by conversion rate. Someone who clicks a citation has already had the basic question answered and is coming to verify or act. Be clear-eyed about the other side of that coin: most people who read an AI answer never click anything, so a citation's value is part high-intent clicks, part being present in the answer your buyer trusts. The same study projects AI search could send sites more visitors than traditional search for digital-marketing and SEO topics by early 2028, though that is an extrapolation, not a measurement. ## How AI Engines Decide What to Cite Google ranks pages. AI engines cite statements. That one difference explains most of what follows. When an assistant answers with sources, it is running some version of retrieval-augmented generation (RAG): fetch relevant documents from a search index, pull the passages that answer the question, synthesize, and attribute. Selection happens at the passage level. The engine does not cite your page because it is good overall. It cites your page because one extractable chunk of it answered one sub-question well. But there is a gate before any of that. The engine has to decide to search the web at all. If it answers from training data alone, nobody gets cited, no matter how optimized the page is. That gate is narrower than most guides admit. In our own analysis of 377 prompts run through ChatGPT's API, citations appeared exactly when the web-search flag fired, and how often it fired tracked intent closely: commercial prompts triggered a live search 72.4% of the time, while purely informational prompts triggered one just 2.5% of the time. A couple of caveats: each prompt ran once (answers vary between runs), and this is one engine's API, whose search routing is not necessarily the consumer app's. Treat the numbers as directional. The practical read survives both: **citation optimization pays off first on commercial and comparison prompts**, because those are the prompts that send engines looking for sources.
![Four stages of how an AI answer selects and links its cited sources.](/blog/how-to-get-cited-by-ai/ai-citation-pipeline.png)
Citations happen at the passage level, and only when the engine searches at all.
## Step Zero: Make Sure AI Can Fetch Your Page Before content, structure, or schema: can the engines' crawlers physically reach your pages? This is the step most guides skip, and the most common failure nobody notices. Each provider runs separate bots for separate jobs, and blocking the wrong one costs you citations without touching rankings. OpenAI [documents its crawlers](https://developers.openai.com/api/docs/bots) separately: **OAI-SearchBot** decides whether you appear in ChatGPT search results, **GPTBot** collects training data, and **ChatGPT-User** fetches pages when a user asks about them live. It is easy to get this wrong in either direction: blocking GPTBot does not remove you from ChatGPT search, and blocking OAI-SearchBot quietly removes you from its generated search answers. Google works the other way: AI Overviews and AI Mode ride on regular Googlebot, so you cannot block the AI features without leaving Search itself. And because ChatGPT's search still leans on Bing's index, blended with OpenAI's own crawl, being indexed in Bing is cheap insurance for the largest assistant: confirm it in Bing Webmaster Tools, where IndexNow can speed fresh pages in. Your CDN can make this decision for you without asking. Cloudflare [now classifies AI crawlers](https://blog.cloudflare.com/content-independence-day-ai-options/) into search, agent, and training bots, and from September 15, 2026, domains newly onboarding to Cloudflare get new defaults: training and agent crawlers blocked on pages that display ads, search crawlers still allowed. Reasonable defaults, but if your firewall rules predate the classification, audit what is actually being blocked rather than assuming. Two more quiet killers: most AI crawlers do not execute JavaScript, so content that only exists after client-side rendering is invisible to them, and paywalled or login-gated pages are effectively out of the running. Our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows which of 34 AI bots your robots.txt allows or blocks; the [Agent Readiness scan](https://geotoolbox.ai/tools/ai-readiness) goes further and live-fetches your pages as each bot to catch WAF rules, 403s, and JavaScript walls. Our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) covers the robots.txt specifics. ## Every Engine Pulls from a Different Index There is no single "AI search" to optimize for. Each assistant grounds its answers in a different index, trusts a different mix of sources, and cites a different number of them. Search Engine Land's [analysis of 8,000 AI citations](https://searchengineland.com/how-to-get-cited-by-ai-seo-insights-from-8000-ai-citations-455284) across 57 queries, run on Rankscale data, put numbers on how differently the engines behave. The percentages below come from that dataset and are a snapshot, not physics: these shares move. The grounding-index column and the Claude row, which SEL did not test, draw on the vendors' own documentation and independent reporting.
EngineGrounding indexWhat it favorsOne move that matters
ChatGPTBing plus OpenAI's own crawlAuthority sources: Wikipedia alone took 27% of its citations in the SEL dataset, news another ~27%, almost no forumsGet documented in neutral reference material, not just your own blog
Google AI Overviews / AI ModeGoogle's indexThe broadest mix: blog-style articles ~46%, news ~20%, plus Reddit, YouTube, and LinkedIn; Wikipedia under 1%Standard Google SEO still buys the ticket; deep pages beat homepages
GeminiGoogle's indexBlogs ~39% and news ~26%, with YouTube as its single most-cited domainVideo is a citation asset here, not decoration
PerplexityIts own indexEditorial and expert review sites, recency, and selective community contentFreshness and niche expert coverage over raw domain authority
ClaudeBrave SearchWell-sourced, measured analysis; conservative citation volumeGoogle rankings don't carry over directly; Brave visibility is its own job
The engines also disagree about how many brands belong in an answer. In the same dataset, ChatGPT and AI Overviews named roughly 3 to 4 brands per answer while Perplexity averaged around 13. If you are a mid-tier brand, your first appearances will most likely come from Perplexity, whose longer lists leave room for smaller names. The mix shifts with intent, too. For B2B queries, company sites and vendor blogs earned about 17% of citations in the SEL data; for consumer queries, official company sites dropped under 4% and review sites, YouTube, and communities took over. We keep engine-specific playbooks for each surface: [ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt), [Perplexity](https://geotoolbox.ai/blog/perplexity-seo), [Gemini](https://geotoolbox.ai/blog/gemini-seo), and [Google AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo). ## Structure Content so a Machine Can Lift It Google's official position on optimizing for AI features is blunt: there are "no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary," per its [AI features documentation](https://developers.google.com/search/docs/appearance/ai-features). Read that carefully. It rules out tricks. It does not rule out clarity, and clarity is what passage-level retrieval rewards. Because engines cite statements, the unit of optimization is the passage, not the page. The moves that consistently show up in cited content: ### Lead Each Section with the Answer A direct, self-contained answer in the first sentences of a section survives extraction. An answer that only makes sense after three paragraphs of setup does not. ### Write Sections That Stand Alone If a chunk needs the surrounding context to be understood, an engine cannot quote it cleanly. Headings that say what the section answers ("Which bots do I allow in robots.txt," not "Key considerations") do half of this work. ### Make Claims with Evidence Attached Specific numbers with named sources give an engine something checkable to attribute. Vague claims give it nothing to cite. ### Cover the Question's Neighborhood on One Page Engines fan a prompt out into implied sub-questions and pull passages per sub-question, so a page that answers the follow-ups too gets more chances to be cited. We wrote up the mechanics in our [query fan-out](https://geotoolbox.ai/blog/query-fan-out) guide. One warning: cover the neighborhood on one deep page. Spinning each sub-question into its own thin URL is exactly what Google's scaled-content policies exist to catch. None of this is exotic. It is the difference between writing to be read and writing to be quoted, and it happens to make pages better for people as well. ## The Myth Pile: llms.txt, Magic Schema, and Recycled Stats Citation optimization has collected a set of tactics that sound technical, cost real hours, and have no evidence behind them. Three are worth naming. **llms.txt.** The proposal is a markdown index that tells AI crawlers what matters on your site. The problem: no engine documents it as a citation signal. Google's own documentation says you do not "need to create new machine readable files, AI text files, or markup" for its AI features. We have taken the same position for our own site: skip it until an engine actually commits, and spend the hour on reachability instead. **Special schema for AI.** The same [Google document](https://developers.google.com/search/docs/appearance/ai-features) is explicit that "there's also no special schema.org structured data that you need to add." Schema is still worth having for what it always did, entity disambiguation and rich results. But treating a markup type as a citation switch does not survive contact with data. In our own 5,234-page citation dataset, HowTo markup was exactly as common on the pages AI engines never cited as on the pages they did - a single-pass reading, but not what you'd expect if it were a citation switch. The honest summary: schema describes your content to machines; it does not make weak content citable. **Recycled statistics.** A surprising share of generative engine optimization (GEO) advice rests on numbers with no traceable primary source. Percentages like "structured content earns 3x more citations" circulate from blog to blog, each citing the previous one, with no methodology in sight. Before you rebuild your content strategy around a statistic, ask who measured it, on what sample, and when. The studies we cite in this article publish their counts, and where we use our own unpublished data we give you the sample and the single-pass caveat so you can weight it accordingly. That filter eliminates most of the genre. ## Earn the Third-Party Mentions Engines Trust Here is the uncomfortable part: much of what determines whether AI cites you does not live on your site. Engines validate brands against the sources they already trust, and for most queries those are third-party surfaces. In the [SEL citation data](https://searchengineland.com/how-to-get-cited-by-ai-seo-insights-from-8000-ai-citations-455284), consumer-intent answers drew on review sites, YouTube, communities, and mainstream press while official company sites took under 4% of citations. Your buyers' assistants are reading about you, mostly not from you. What that means in practice depends on your market. For B2B, being present in industry publications, comparison articles, and directories is what puts you in the consideration set the engines draw from; in the SEL data that meant niche trade outlets like TechTarget, directories like Clutch, analyst reports from Gartner and Statista, and LinkedIn expert posts. For consumer brands, reviews and community threads carry more weight, and the Google-side engines have a real appetite for user content: Reddit and Quora took 2 to 5% of their citations in the SEL data, against under 0.5% for ChatGPT. A well-maintained entity footprint (a consistent description of what you are on Wikipedia or Wikidata where warranted, LinkedIn, Crunchbase, review profiles) gives the engines a consistent description to draw on instead of leaving them to improvise your positioning. What keeps this from sliding into spam: first, participate where you can add something real; engines and communities both punish manufactured advocacy, and seeded threads have a way of becoming the story. And second, the strongest third-party play is still publishing something worth citing: original data, a benchmark, a methodology. Give other sites a reason to reference you, and the engines inherit that judgment. ## Track Whether You're Actually Getting Cited A single check proves nothing. SparkToro and Gumshoe [had 600 volunteers run prompts 2,961 times](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/) across ChatGPT, Claude, and Google's AI, and found under a 1-in-100 chance that ChatGPT or Google's AI returns the same brand list in any two runs of the same prompt. Checking once and concluding you are visible (or invisible) is reading noise. The metric they land on as usable is visibility rate: the share of many runs, across many prompts, where you appear.
What to trackFree methodWhere it falls short
Citation rate (your pages linked as sources)Run 10-20 buyer prompts in each engine, several times each, and log linked sources per runManual, and personalization can skew what you see
Mention rate (your brand named, linked or not)Same prompt runs, logging brand names in the answer textMisses how you are described unless you record wording
AI referral trafficA GA4 channel group matching chatgpt.com, perplexity.ai, claude.ai, gemini.google.com referrersA starter list: app traffic and stripped referrers slip through, and AI Overviews clicks hide inside normal Google traffic
Citation lossRe-run your prompt set monthly and diff which sources appearCited source sets shift between months, so confirm a loss before reacting to it
The free method works and we documented the full protocol, including how many runs you need before trusting a result, in our guide to [tracking brand mentions in AI search](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search). It stops scaling once you are past a few dozen prompts and more than a couple of engines. That is where geotoolbox's [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) takes over: it captures which pages, yours and your competitors', each engine cites for your prompt set, so a lost citation shows up in your next scan instead of going unnoticed. Like any tracker, it samples, so the same run-count discipline applies to its numbers too. ## A 30-Day Plan to Earn Your First Citations You do not need a quarter-long program to find out whether this works for your site. One month, four moves: **Week 1: verify reachability.** Test your key pages against each engine's bots, fix robots.txt and CDN rules, and confirm your money pages render without JavaScript. This is an afternoon of work that everything else depends on. **Week 2: baseline.** Write 10-20 prompts your buyers would really ask, weighted toward commercial and comparison intent. Run each several times per engine and log mentions, citations, and who gets cited instead of you. The competitors' cited pages are your reverse-engineering material. **Week 3: restructure three pages.** Pick the three pages closest to the prompts where competitors get cited and you do not. Give every section an answer-first opening, headings that state the question, and claims with sourced numbers. Do not create new thin pages; deepen the ones you have. **Week 4: one asset, two placements.** Publish one piece of original data, even a small one (a survey of your customers, a benchmark from your product's usage), and pitch it to two industry publications or communities where your buyers already look. Then re-run the Week 2 prompt set and compare. Citations compound slowly, and a month will not make you ChatGPT's favorite source. It will tell you where you stand, clear the blockers you could not see, and put the first evidence-backed passages in front of the engines. ## Frequently Asked Questions ### How long does it take to get cited by AI? There is no reliable published benchmark, and anyone quoting a precise timeline is guessing. What is knowable: engines with live retrieval (Perplexity, ChatGPT search, AI Overviews) can cite a page once their underlying search engine indexes it, so reachability and indexing set the floor. Indexing only makes you eligible, though: retrieval and selection still decide. Authority-driven citation, being the source engines prefer, builds over months of third-party presence, not days. ### Do you need a Wikipedia page to get cited by AI? No, and for most businesses pursuing one is wasted effort. Wikipedia dominates ChatGPT's citations for encyclopedic queries, but commercial and comparison prompts draw from review sites, editorial coverage, and vendor content instead. If your brand does not meet Wikipedia's notability bar, put the energy into the sources engines cite for buying questions. ### Do AI citations help your Google rankings? Not directly; there is no evidence citations feed back into ranking algorithms. The relationship mostly runs the other way: for Google's AI surfaces, ranking well makes citations more likely. The overlap is in the inputs, since the same clarity, evidence, and third-party authority that earn citations also support rankings. ### Which AI assistants actually show citations? Perplexity attaches sources to virtually every search answer. Google AI Overviews and AI Mode link sources whenever they appear. ChatGPT cites when it searches the web, which depends on the prompt. Gemini and Claude cite when their search grounding fires. If you are choosing where to start measuring, start where citations are most consistent: Perplexity and Google's AI surfaces. ### Is llms.txt worth creating? No engine documents llms.txt as a citation signal, and Google has said outright that its AI features need no new files, so we skip it on our own site. If you already have one, leaving it up costs you nothing but upkeep; just do not mistake maintaining it for progress. ### What should you do if you lose a citation? Run the diagnostic in order: confirm the loss is real (re-check the prompt several times), then reachability (did a CDN or robots rule change?), then the replacement (is the page that took your slot fresher or more specific?), then update yours accordingly. Source sets change frequently, which cuts both ways: losses are common, and so are second chances. ## Get Cited, Then Stay Cited So, how do you get cited by AI? Make your pages fetchable by every engine's bots, write passages a machine can lift whole, back claims with numbers worth quoting, and build the third-party footprint engines check you against. Then measure like the answers are nondeterministic, because they are. The brands winning citations right now are not running secret tactics. They are the ones who did the unglamorous parts first and then watched the engines closely enough to notice what changed. If you would rather not do the watching by hand, our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) runs your prompt set for you and shows exactly which pages each engine cites, yours and your competitors', so gains and losses turn up in data instead of anecdotes. ## Sources - How to get cited by AI: SEO insights from 8,000 AI citations - Search Engine Land, 2025 - `searchengineland.com/how-to-get-cited-by-ai-seo-insights-from-8000-ai-citations-455284` - AI features and your website - Google Search Central documentation, updated December 2025 - `developers.google.com/search/docs/appearance/ai-features` - OpenAI crawlers and bots documentation - OpenAI - `developers.openai.com/api/docs/bots` - Your site, your rules: new AI traffic options for all customers - Cloudflare, 2026 - `blog.cloudflare.com/content-independence-day-ai-options/` - AIs are highly inconsistent when recommending brands - SparkToro / Gumshoe.ai, 2026 - `sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/` - Semrush AI search traffic study - Semrush, 2025 - `semrush.com/blog/ai-search-seo-traffic-study/` --- ## What Is Qwen3.8-Max? Alibaba's 2.4T Flagship > What is Qwen3.8-Max? Alibaba's 2.4T-parameter flagship: the published benchmarks, the independent ones, API pricing, and the open-weights promise. - Canonical: https://geotoolbox.ai/blog/qwen3-8-max - Published: 2026-07-22 · Updated: 2026-08-18 When Alibaba previewed Qwen3.8-Max in July, half the AI assistants we tested did not know it existed, and one that did was already repeating specs Alibaba had never published. Two weeks later the real release landed with a full benchmark table, and on its headline coding benchmark the independent number came in five points below Alibaba's own. This page pins down what is actually known as of August 14, 2026: what the published benchmarks show and where the outside measurements disagree, what the API costs, what the open-weight release did and did not deliver, and why a model most marketers will never touch still shapes what AI tools say about their brands. Every number below is labeled as Alibaba's own or independently measured. ## What Is Qwen3.8-Max? **Qwen3.8-Max is Alibaba's current flagship AI model, released on August 3, 2026: a 2.4-trillion-parameter multimodal system with 95 billion parameters active per token, a one-million-token context window, and a price of $2 per million input tokens and $6 per million output.** It takes text and images natively, with Alibaba also claiming video and document handling, and you call it as `qwen3.8-max` through Alibaba Cloud Model Studio. Alibaba calls it ["the most capable model in the Qwen family to date"](https://qwen.ai/blog?id=qwen3.8) and says it improves on Qwen3.7-Max across coding, work, research, and long-horizon tasks. That is a notably quieter claim than the one that trailed the July preview, when [Alibaba's internal evaluations](https://www.scmp.com/tech/article/3361119/alibaba-says-newest-qwen-ai-model-second-only-anthropics-claude-fable-5) had it "second only" to Anthropic's Claude Fable 5. The release post drops that line, and the benchmark table it published instead is the reason why. One thing to establish before reading those benchmarks, because it decides which rows mean anything. On some rows Alibaba re-ran every model itself: its SWE-bench Pro footnote says all baselines were evaluated on the same corrected benchmark. On its headline row, Terminal-Bench 2.1, it did not. There it ran Qwen under Claude Code and, in its own words, reported "the best published score across harnesses" for everyone else, taking Claude's figures from Artificial Analysis and GPT-5.6 Sol's from OpenAI. A self-run score against other labs' best published scores is not a controlled comparison. The competitive timing was not subtle. The July preview landed days after Moonshot AI announced [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3), the model Moonshot's own card puts at 2.8T parameters, which had just claimed the "largest ever" headline and then shipped its weights on schedule. Alibaba then promised something it had never done: open weights for a Max-class model, committed for the week of August 10, 2026. It delivered two days past that window, on August 12, with a text-only variant under a bespoke license. For a company whose [Qwen family](https://geotoolbox.ai/blog/what-is-qwen) built its reputation on open models, what that release does and does not include matters, and we get to it below. One housekeeping note: Alibaba's Qwen team also shipped Qwen-Image-3.0 on July 21. That is a separate image-generation model, despite the similar naming and back-to-back launch dates. ## Published vs Independently Confirmed: The Spec Reality The July preview shipped with almost nothing you could check. The August release fixed most of that, and the distinction that still matters is between what Alibaba has now published and what anyone outside Alibaba has confirmed. The table below shows where each spec stands as of August 14, 2026.
SpecStatusWhat we actually know
2.4T total parameters, 95B activeVendor-publishedBoth figures are now in Alibaba's own release post; the July preview disclosed only the total
Sparse MoE architectureVendor-publishedConfirmed by the 95B-active figure, though Alibaba has still not published the full configuration or a technical report
Multimodal (text, image, video, docs)Partly publishedAlibaba's model config lists only text and image as input modalities; video handling is evidenced by its own video benchmark rows, and document handling is still its description rather than a spec
1M-token context windowVendor-publishedNo longer community lore: Alibaba's model config sets a 1,000,000-token window, and Model Studio bills against a 0 to 1M tier
$2 in / $6 out per 1M tokensVendor-publishedListed on Alibaba Cloud Model Studio's pricing page for the international region, the same rate in thinking and non-thinking modes
Benchmark scoresPublished, and contestedA full table shipped with the release. Some rows re-run every model; the headline Terminal-Bench row imports rivals' best published scores. On that row Artificial Analysis independently measures 81.3% where Alibaba reports 86.6
"Second only to Fable 5"Withdrawn in practiceThe August release post does not repeat it, and neither Alibaba's own table nor the independent leaderboard supports it
Open-weight releaseDelivered August 12, 2026 (partial)Qwen3.8-2.4T-A95B on Alibaba's official Hugging Face org: text-only, 262,144-token native context (extensible to about 1M), bespoke "Qwen3.8-Max" license, not Apache 2.0
The single most useful addition is the active-parameter count. In a [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) model only a fraction of the total parameters fire on each token, and that fraction drives compute cost per token while the 2.4T total sets the memory needed just to hold the weights. With 95B active, Qwen3.8-Max activates about 4% of itself per token, which keeps per-token compute far below what the total size suggests and makes a $2 input price economically plausible. In July that number was missing and serving economics were guesswork. For the generational comparison, Alibaba's table carries a Qwen3.7-Max column on every row: SWE-bench Pro goes 60.6 to 67.7, GPQA Diamond 92.4 to 92.6, and DeepSWE 1.1 jumps 21.6 to 56.6. Treat the Terminal-Bench row differently, because there the 3.7-Max figure is imported from Artificial Analysis while the 3.8-Max figure is Alibaba's own, so 74.5 to 86.6 is not one measurement against another. Measured by Artificial Analysis at both ends, the generational gain on that benchmark is 74.5 to 81.3. Separately, the older [Qwen3.7 benchmark card](https://www.alibabacloud.com/blog/qwen3-7-the-agent-frontier_603154) reports 80.4% on SWE-bench Verified, a different test from the new table's SWE-bench Pro, so those two coding numbers do not belong on the same axis either. ## How Do You Access Qwen3.8-Max? The ordinary way in is now the API. Qwen3.8-Max is a standard pay-per-token model on Alibaba Cloud Model Studio, called as `qwen3.8-max`, at $2 per million input tokens and $6 per million output in the international region. That rate is flat across the full million-token context and applies in both thinking and non-thinking modes, with a context-caching discount on repeated input and a one-million-token free allowance to get started. This is the biggest practical change since the July preview, which had no per-token price at all. The subscription route still exists alongside it. Token Plan sells credit bundles in Lite, Standard, and Pro tiers, and Qoder and QoderWork are Alibaba's agentic coding products. Those made sense as the only door during the preview, when [a limited-time campaign](https://docs.qoder.com/events/qwen-max-preview) ran the preview build at up to 90% off its standard credit coefficient, and 98% off overnight in Singapore hours. Treat that campaign as preview-era: it was attached to `qwen3.8-max-preview`, and we have not re-verified it against the released model. Price the API rate first and check any credit promotion live before you count on it. One caution stands regardless of which door you use. Know who you are contracting with: the international Token Plan checkout names Intelligent Cloud Computing (Singapore) Private Limited as the operator rather than an alibabacloud.com property. That is consistent with Alibaba running international billing through a Singapore entity, but confirm the contracting entity and terms against your own Alibaba Cloud account before entering payment details or API keys. For developers, the friendliest detail is protocol compatibility: the endpoint speaks both the OpenAI and Anthropic API formats, so it drops into Claude Code, Cursor, OpenCode, and similar tools with a config change. It also exposes a `reasoning_effort` setting with three levels, `xhigh`, `medium`, and `low`, which is the lever for trading cost against depth. Alibaba ships `xhigh` as the default, so the out-of-the-box configuration is the most expensive one. For context on how rivals price this tier, our [Kimi API pricing breakdown](https://geotoolbox.ai/blog/kimi-api-pricing) is the comparison point: $3 per million input tokens and $15 per million output, against Qwen's $2 and $6. ## How Does It Compare to Kimi K3 and Fable 5? It depends on who is holding the stopwatch, and that turns out to be the story. Start with Alibaba's own table. Against Claude Fable 5 it reports Qwen3.8-Max ahead on Terminal-Bench 2.1, 86.6 to 84.6, and level on GPQA Diamond at 92.6. It also reports Fable 5 comfortably ahead on the harder tests: 80.0 to 67.7 on SWE-bench Pro, and 53.3 to 43.6 on Humanity's Last Exam. Against GPT-5.6 Sol the split runs the other way again, with Qwen behind on Terminal-Bench 2.1 and GPQA Diamond but ahead on SWE-bench Pro, 67.7 to 64.6. Then check the same benchmark somewhere Alibaba does not control. [Artificial Analysis independently runs Terminal-Bench 2.1](https://artificialanalysis.ai/evaluations/terminalbench-v2-1) and scored Qwen3.8-Max at **81.3%** as of August 7, 2026, which puts it *below* Claude Fable 5's 84.6% and Kimi K3's 85.0% on that same leaderboard, and tenth overall. Alibaba reports 86.6% for the same model on the same benchmark. Alibaba disclosed the method that produces the gap: it ran Qwen under Claude Code at avg@10 with a five-hour timeout, and for every other model reported the best published score across harnesses. So the two 86.6 and 81.3 figures are the same model on the same benchmark measured by different people under different settings, and a five-point spread there is ordinary rather than scandalous. Neither party has published the reasoning-effort level it ran at, and Qwen3.8-Max defaults to its most expensive setting, which is exactly the kind of variable that moves an agentic score. What it does mean is narrower, and still worth knowing. Under the only harness anyone outside Alibaba has run on the released model, the ordering on Alibaba's headline benchmark flips. There is one more outside test, and it predates the release. [Trilogy AI's StackPerf run](https://trilogyai.substack.com/p/qwen-38-max-benchmark-how-it-compares) is a blind architecture-analysis benchmark where Qwen3.8-Max-Preview and Kimi K3 each analyzed an identical 269-file codebase under identical limits, with the reports scored blind. Kimi K3 scored 83 out of 100; Qwen3.8-Max-Preview scored 80. Note the build: that was the preview, not the released model, so it describes the July snapshot rather than what ships today. The same run caught a real Qwen strength, 22 gateway requests and 44 tool calls against Kimi's 53, none of Qwen's failed, and the blind review scored its tool use 9 to Kimi's 8. One thing the coding numbers understate: Alibaba published a second table for multimodal work, and Qwen3.8-Max wins most of it. It leads Fable 5 on MathVision (95.2 to 92.7), LogicVista (91.9 to 85.7), HiPhO (90.0 to 78.6) and most of the vision rows, while the video rows split roughly evenly between them. Those are still vendor-run, and no one has independently checked them, but if you are choosing a model for document, image, or video work rather than for terminal coding, the picture is considerably better than the headline benchmark suggests.
ModelParametersBenchmark positionPricing per 1M tokensWeights
Qwen3.8-Max2.4T total, 95B active (vendor-stated)Vendor table: 86.6 Terminal-Bench 2.1, 67.7 SWE-bench Pro, 92.6 GPQA. Independently, Artificial Analysis measures 81.3 on Terminal-Bench 2.1$2 in / $6 outOpen since August 12, 2026 (text-only variant, bespoke "Qwen3.8-Max" license); multimodal version API-only
Kimi K32.8T (claimed)StackPerf 83/100; AA Intelligence Index 57$3 in / $15 outPublished on Hugging Face under a custom Kimi K3 License
Qwen3.7-Max (previous flagship)UndisclosedVendor card: GPQA 92.4, SWE-bench Verified 80.4%$2.50 in / $7.50 out list, on a limited-time 50% discountClosed
Claude Fable 5UndisclosedSWE-bench Verified 95% (vendor)$10 in / $50 outClosed
The benchmark names in that table are not interchangeable. SWE-bench Pro, SWE-bench Verified, and StackPerf are three different tests run by three different parties, so each row only means something on its own terms. Do not compare a score in one row against a score in another. Early hands-on testing shows what the benchmark cannot. [One reviewer](https://thomas-wiegold.com/blog/qwen-3-8-max-review/) who runs the same four build prompts on every flagship found the July preview build one-shotted a full Texas Hold'em simulation in Go, something only Claude Fable 5 and Grok 4.5 had managed on his suite, and produced what he called possibly his best-ever result on a website build. The cost: it was the slowest model he had ever tested, taking 1 hour 20 minutes on the poker task, with long stretches of thinking. If the preview exposed the same `xhigh` default the released endpoint ships with, that would account for much of it. His testing predates the August release either way, so read the slowness as preview-era; the build quality he describes is still the most useful hands-on signal anyone has published. The gap at the top is still real, and the cleanest like-for-like is Alibaba's own SWE-bench Pro row: Fable 5 at 80.0 against Qwen3.8-Max at 67.7. That row is the most citable of the ones Alibaba published, for two reasons: its footnote says every baseline was re-run on the same corrected benchmark rather than imported, and it is a vendor number that runs against the vendor's own interest. Kimi K3, [GLM-5.2](https://geotoolbox.ai/blog/chinese-ai-models-compared), and MiniMax all chased the frontier this year and fell short, and Alibaba's own numbers say it did not close the gap either. The useful reading is narrower than the July headline: Qwen3.8-Max is a strong multimodal model and a competent coding one, priced far below the frontier, and nothing yet published establishes it as the second-best model in the world. One deployment note for regulated businesses: like the rest of the family, Qwen models carry documented, baked-in guardrails aligned with Chinese content rules on politically sensitive topics. For coding and agent work this rarely surfaces; for content-adjacent or compliance-sensitive workloads, test before you commit. Data residency deserves the same diligence: confirm the service region, retention terms, and contracting entity for the specific product rather than assuming them. If residency is a hard requirement, the clean answer is open weights and self-hosting. ## Did the Weights Open? The open-weight promise was the biggest claim in the release and the one with the weakest track record behind it, and it was ultimately kept, two days late and with caveats. **On August 12, 2026, [Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) landed in Alibaba's official Hugging Face organization.** The published variant is text-only (the multimodal version stays API-only), its native context stops at 262,144 tokens (extensible to about 1M), and the license is a bespoke "Qwen3.8-Max" license rather than Apache 2.0. Here is the pattern it broke. Alibaba runs a genuine two-track strategy: the small and mid-size Qwen models ship under Apache 2.0 and dominate the open ecosystem, while the Max-tier flagships stayed closed. Qwen3.7-Max: API-only. Qwen3.6-Max-Preview: API-only. No Max-tier flagship had ever had its weights published, which Alibaba itself acknowledged by calling this "the first time we will open-source the weights of a Qwen-Max-class model." The repository now exists; what remains open is how the bespoke license terms compare to Apache 2.0 in practice. The competitive read is hard to ignore: Moonshot announced Kimi K3 with a firm weights date of July 27, days before Alibaba's preview, and then met it. The promise kept Alibaba's open-flagship credibility alive through a news cycle it would otherwise have lost, and weight-release promises across the industry have a habit of slipping once headlines move on. The verification bar is simple, and worth being strict about: an actual repository under [Alibaba's own Hugging Face organization](https://huggingface.co/Qwen) containing weights and a license file. Not a blog post, not a tweet, and not a third-party repository that merely carries the name. Repos titled Qwen3.8 have already appeared from unaffiliated accounts, and they are worth understanding because they are the trap. Some are empty placeholders holding a README. Others do contain real weights that are simply not this model: one of the most-downloaded, at roughly 23,000 downloads, ships files named `qwen3-4b-thinking-2507`, an older and far smaller Qwen wearing the new number. Both kinds carry an Apache 2.0 label that is not Alibaba's terms. Until the real URL exists under Alibaba's own account, treat the release as a roadmap item that can slip. And keep in mind that even a delivered release can be less open than the headline suggests: [open weights are not open source](https://geotoolbox.ai/blog/open-weights-vs-open-source), and the license attached will matter as much as the download link. ## Can You Actually Run It? Not this one, and probably not for a long time, even if the weights land tomorrow. At 2.4 trillion parameters, the arithmetic is brutal. The packed weights alone would occupy roughly 1.2 terabytes at 4-bit quantization, before runtime overhead and caches, and even an extreme 1.5 bits per weight leaves hundreds of gigabytes of raw weights, a point Hacker News commenters worked out within hours of the announcement. Serving it properly means a multi-accelerator cluster; a single Nvidia H200 carries 141 GB, so even an eight-GPU node cannot hold the weights. That math works for inference providers and almost nobody else. The 95B active-parameter figure changes the cost picture without changing that conclusion. Compute per token is closer to a 95-billion-parameter dense model than to a 2.4-trillion one, which is why a $2 input price is not as implausible as the headline size makes it look, but you still have to hold all 2.4 trillion parameters in memory to serve any of them. Sparse activation makes the model cheap to run at scale, not cheap to own. The realistic play for anyone who [runs models locally](https://geotoolbox.ai/blog/run-llm-locally) has always been to look at the family, not the flagship, and that promise was kept: **[Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B) shipped on Hugging Face around August 14, 2026**, a 27-billion-parameter native vision-language model, under Apache 2.0, that a single well-specified machine can actually run. Where the Max flagship's own multimodal weights stay API-only, this smaller model is both multimodal and genuinely open, which makes it the more useful release of the two for local and self-hosted work. The 2.4T Max model itself, open or not, remains a datacenter artifact. ## What AI Engines Say About Qwen3.8-Max We at geotoolbox ran a four-engine check on July 22, 2026, three days after the preview and two weeks before the model was actually released, asking each assistant what Qwen3.8-Max is. The split was total. What follows is a snapshot of that launch window, not a description of what the engines say today, and the launch window is exactly the interesting part.
![Panel testing four AI engines on Qwen3.8-Max three days after its preview launch: ChatGPT with web access gave an accurate hedged summary, Perplexity with web access stated unpublished specs as fact, while Gemini and Claude without web access did not know the model existed.](/blog/qwen3-8-max/ai-engine-awareness-qwen3-8-max.png)
In this four-engine check, three days after the preview launch, recognition tracked web access rather than model quality.
Both web-connected engines recognized the model. ChatGPT with browsing summarized the launch accurately from news coverage and correctly hedged the unverified specs. Perplexity knew it too, but went further than the record supports: it asserted the sparse-MoE architecture and a specific context window as fact, sourcing them from third-party API-router pages rather than anything Alibaba published. The two engines answering from training data alone were blank. Gemini suggested the name might be "a typo or confusion," guessing we meant a different model entirely, and described Qwen2.5 as current. Claude declined to guess at all, saying it would not invent specifications for a model it could not verify. The Perplexity result is the one worth dwelling on. For a three-day-old model the blanks are expected behavior; the overclaim is not. In a fresh news window, retrieval-based engines repeat whatever the handful of live pages say, including specs no primary source has confirmed. The unofficial one-million-token context figure it asserted came from an API reseller's page, not from Alibaba. Alibaba's own documentation later confirmed that number, which is the luckiest possible outcome and not the point: at the time it was asserted, nothing supported it, and a guess that happens to be right is still a guess. This is how unverified claims fossilize: the early pages get cited, the citations get repeated, and by the time official documentation lands, the answer engines have weeks of momentum behind whatever the first pages said. The citation data shows where those answers come from. In DataForSEO's LLM-mention tracking for Qwen topics, the most-cited domains are YouTube, Reddit, Hugging Face, and GitHub, community surfaces rather than Alibaba's own properties. Whoever publishes the clearest early page, accurate or not, tends to become the reference. ## What Qwen3.8-Max Means for Your AI Visibility If you are not shipping AI models, here is why this launch belongs on your radar: every new frontier model is a new surface that describes brands, including yours, to its users. Qwen matters disproportionately here because of how it spreads. As we covered in our [Qwen explainer](https://geotoolbox.ai/blog/what-is-qwen), it is the most-downloaded open model family in the world, fine-tuned and rebranded inside products that never mention Alibaba. Now that both the Max weights and the smaller, multimodal Qwen3.8-27B have shipped, they will propagate into that same ecosystem, and their answers about your company travel with them, into tools you will never audit. The launch-window lesson from our four-engine test applies directly to brands. The moment something new is true about your company, a launch, a rename, a pricing change, there is a window where AI engines only know what a few early pages say. Whoever fills that window shapes the answer, and corrections tend to lag well behind the first version. The time to know how AI systems describe your brand is before a wrong version hardens, not after a prospect quotes it back to you. That check takes minutes. Our free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) shows whether AI crawlers can actually reach and read your site, one of the mechanical layers under retrieval-based answers about you, and geotoolbox tracks what the major engines are saying once you are visible. New models will keep launching on someone else's schedule. Your brand's answer should not depend on which one a customer happens to ask. ## Frequently Asked Questions ### Is Qwen3.8-Max released yet? Yes. Alibaba released it on August 3, 2026, after showing a preview build on July 19. It is generally available on Alibaba Cloud Model Studio as `qwen3.8-max`, with published specifications, a benchmark table, and standard per-token pricing. The separate `qwen3.8-max-preview` build was the July stopgap and is superseded. ### Is Qwen3.8-Max free or open source? The API is paid at $2 per million input tokens and $6 per million output, with a one-million-token free allowance to start. The weights opened on August 12, 2026, the first open Max-tier Qwen ever, but with conditions: the published variant is text-only and ships under a bespoke "Qwen3.8-Max" license rather than Apache 2.0, so it is open-weight, not open source. ### How many parameters does Qwen3.8-Max have? 2.4 trillion total, with 95 billion active per token in a sparse mixture-of-experts design. That puts it among the largest publicly disclosed parameter counts anywhere, just below the 2.8T Moonshot states for Kimi K3. Both figures come from Alibaba's own release post; there is still no independent technical report. ### Is Qwen3.8-Max better than Kimi K3? On the evidence available, Kimi K3 is ahead. Artificial Analysis, which measures both under the same harness, scores Kimi K3 at 85.0% and Qwen3.8-Max at 81.3% on Terminal-Bench 2.1. Trilogy AI's earlier blind StackPerf run pointed the same way, 83 to 80, though that one tested the July preview build and Qwen used fewer tool calls with none failing. Alibaba's own published table does not compare the two directly. ### When did the open weights come out? August 12, 2026, two days past the committed week-of-August-10 window: the Qwen3.8-2.4T-A95B repository under Alibaba's own Hugging Face organization, with weights and a bespoke "Qwen3.8-Max" license file. That was the credible bar; the third-party repos that carried the Qwen3.8 name before that date held either nothing or relabelled weights from older Qwen models. For contrast, Moonshot named July 27 for Kimi K3's weights at announcement and delivered on it exactly. ### How can I try Qwen3.8-Max from the US? Call the API directly on Alibaba Cloud Model Studio at $2 per million input tokens and $6 per million output; the endpoint speaks both the OpenAI and Anthropic formats, so it drops into most existing tooling. Token Plan, Qoder, and QoderWork remain available as credit-based alternatives. The international checkout accepts US customers, though some early users report payment friction, so verify the contracting entity against your Alibaba Cloud account first. ## Sources - Qwen3.8-Max: A New Bar for Coding and Cowork (official launch post, specs, benchmark table, open-weights commitment) - Qwen / Alibaba, August 3, 2026 - `qwen.ai/blog?id=qwen3.8` - Alibaba Cloud Model Studio model pricing (official, qwen3.8-max at $2 in / $6 out) - Alibaba Cloud - `alibabacloud.com/help/en/model-studio/model-pricing` - Text generation models, Model Studio (official API model id) - Alibaba Cloud - `alibabacloud.com/help/en/model-studio/text-generation` - Terminal-Bench v2.1 leaderboard (independent measurement, Qwen3.8-Max at 81.3%) - Artificial Analysis - `artificialanalysis.ai/evaluations/terminalbench-v2-1` - Qwen models - Hugging Face (checked for the open-weight release, August 7, 2026) - `huggingface.co/Qwen` - Alibaba Previews Qwen3.8-Max, a 2.4 Trillion-Parameter Multimodal Model - MarkTechPost, July 19, 2026 - `marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch` - Alibaba says newest Qwen AI model is second only to Anthropic's Claude Fable 5 - South China Morning Post, July 19, 2026 - `scmp.com/tech/article/3361119/alibaba-says-newest-qwen-ai-model-second-only-anthropics-claude-fable-5` - Qwen3.8-Max-Preview limited-time benefit (official promotion doc) - Qoder Docs, July 2026 - `docs.qoder.com/events/qwen-max-preview` - Qwen 3.8 Max Benchmark: How It Compares With Kimi K3 - Trilogy AI, July 2026 - `trilogyai.substack.com/p/qwen-38-max-benchmark-how-it-compares` - Qwen3.8-Max Review - thomas-wiegold.com, July 2026 - `thomas-wiegold.com/blog/qwen-3-8-max-review` - Qwen 3.8 (discussion thread, 959 points) - Hacker News, July 2026 - `news.ycombinator.com/item?id=48966120` --- ## How to Run an LLM Locally (2026 Guide) > Run a large language model on your own Mac or PC. A practical guide to hardware, tools, quantization, setup, using it for coding, and whether it's worth it. - Canonical: https://geotoolbox.ai/blog/run-llm-locally - Published: 2026-07-22 · Updated: 2026-07-28 Running an LLM locally used to mean a weekend of compiling code and reading GPU forums. In 2026 it takes one command and about fifteen minutes, and it works on the laptop you already own. This guide covers the parts that actually trip people up: whether your hardware can handle it, which tool to install, what those cryptic model file names mean, how to wire a local model into your code editor, and the honest answer to whether any of it is worth the effort. It is for developers and marketers who want a private, offline model of their own, without a per-token bill. You do not need a background in machine learning, just a willingness to run a command or so. ## What Running an LLM Locally Actually Means A local LLM is an open-weight language model that runs entirely on your own computer instead of calling a hosted service like ChatGPT or Claude. You download the model file once, and from then on every prompt is answered by your own hardware. Nothing is sent to a vendor, nothing is logged on someone else's server, and it keeps working with the Wi-Fi off. That is the whole appeal, and also the whole catch. A model running on your laptop is not a free copy of ChatGPT. It is a different tool that happens to share the same interface. It excels at private, repetitive, and offline work, and it is genuinely useful for learning [how these models actually work](https://geotoolbox.ai/blog/how-ai-models-think). It will not match a frontier cloud model on the hardest reasoning or on very long documents, and the largest open-weight models are a different exercise entirely: see [how to run Kimi K3 locally](https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally) for what self-hosting a 2.8-trillion-parameter model actually takes. Set that expectation early and everything that follows makes sense. Get it wrong, and you will spend an evening fighting your GPU only to conclude that local models are "dumb." They are not dumb. They are smaller, and you are the one deciding how small. You have a few decisions to make: what your machine can handle, which tool to install, and which model file to download. The rest of this guide walks through them in order. ## Can Your Computer Run One? Memory is the gate. Before download size or GPU speed, the question is whether the model fits, because a model that fits in fast memory runs many times quicker than one that spills over into slower memory. Three numbers matter, and only one applies to you: - **VRAM** is the memory on a discrete graphics card (an NVIDIA or AMD GPU). On a Windows or Linux desktop, this is the number that counts. - **System RAM** is your main memory. A model can run here on the CPU alone, but slowly. - **Unified memory** is Apple Silicon's trick: on an M-series Mac, the CPU and GPU share one pool, so a Mac with plenty of unified memory can load models that would need an expensive discrete GPU on a PC. This is why a mid-range Mac is often the smoothest place to start. A rough rule at the common Q4_K_M quantization (explained below): budget roughly half a gigabyte for the weights of each billion parameters, then add headroom for the runtime and your conversation. A 7-billion-parameter model needs roughly 5 GB all in, so an 8 GB card handles it with room to spare; an 8B model fits too, with less headroom for long chats.
Your memoryLargest model (Q4_K_M)What to expect
8 GB RAM or VRAM7B to 8BRoughly 30 to 50 tokens/sec on a GPU; single digits, and slow, on CPU only
16 GB RAM or VRAM13B to 14BThe everyday sweet spot; roughly 20 to 40 tokens/sec on a GPU
24 GB VRAM (RTX 3090/4090)up to 32BNear the best quality you can run locally; roughly 15 to 30 tokens/sec
48 GB+ (2x24GB, or a large-memory Mac)up to 70BHighest local quality; slower, roughly 8 to 15 tokens/sec, and worth it for the size
Those speeds are rough and assume the model runs on a GPU or Apple Silicon; the exact number depends on your specific chip and the model. Interactive chat is comfortable once you are into the double digits of tokens per second, and the largest models trade speed for quality by design. One counterintuitive point: capacity beats raw speed. An older 24 GB RTX 3090 will out-run a newer 16 GB RTX 4080 on a large model, because the 3090 keeps the whole model in VRAM while the 4080 has to offload part of it to system RAM and pays the speed penalty. No GPU at all? You can still run small models on the CPU, but expect roughly 3 to 8 tokens per second on a 7B model, which is fine for a background task and painful for a live chat. ## Pick Your Tool Half a dozen apps all claim to "run local LLMs," which is where most people stall. Sort them by the job you actually have, not by feature lists. The real split is command line versus graphical interface, and beginner versus builder.
ToolInterfaceBest forWorth knowing
OllamaCommand line + APIDevelopers who want a model behind an APIOne command to install and run; serves an OpenAI-compatible endpoint by default
LM StudioDesktop GUIBeginners who want to point and clickSearches and downloads models in-app; no terminal needed
JanDesktop GUIPrivacy-first usersOpen source, fully offline, a clean ChatGPT-style window
GPT4AllDesktop GUIAbsolute beginners on modest hardwareRuns well CPU-only; chatting in a couple of minutes
llama.cppCommand line / engineTinkerers who want full controlThe engine most other tools are built on, including Ollama
A couple more are worth naming even though they are not starting points. [Open WebUI](https://docs.openwebui.com/) puts a ChatGPT-style browser interface on top of Ollama when you outgrow the terminal. [vLLM](https://github.com/vllm-project/vllm) is a serving engine for when a model needs to handle many users at once: Ollama is tuned for one person on one machine, while vLLM is built to batch many simultaneous requests, so a busy shared server is a different problem from running a model on your own desk. If you are not sure, the short version: install [Ollama](https://docs.ollama.com/quickstart) if you are comfortable in a terminal or want to build on it, and [LM Studio](https://lmstudio.ai/) if you would rather never see one. Both are free, both run the same underlying models, and you can switch later without losing anything. If leaving no trace is the whole point, Jan and GPT4All lean hardest into privacy and are built to run fully offline by default. ## Understand the Model Files (Quantization and GGUF) Go to download a model and you hit a wall of cryptic file names: the same model in versions labeled Q2_K, Q4_K_M, Q5_K_M, Q8_0. This is the step that stops people cold, and it is simpler than it looks. Those labels are **quantization** levels. A model's weights are originally stored at high precision (16 bits each). Quantization rounds them to fewer bits to shrink the file and speed up inference. Four-bit quantization cuts the memory footprint by roughly 75 percent while losing only a small amount of quality, usually in the low single digits of a percent. That trade is what makes an 8-billion-parameter model fit on a laptop at all. **GGUF** is the standard file format these quantized models come in, used by Ollama, LM Studio, llama.cpp, and Jan. When you grab a model from [Hugging Face](https://huggingface.co/), you are looking for its GGUF version. The default you want, in almost every case, is **Q4_K_M**. It is the accepted sweet spot: small enough to fit, good enough that you will not notice the compression. Step up to Q5_K_M or Q6_K if you have memory to spare and want a little more quality. Avoid Q2 and Q3 unless you are desperate to squeeze a model onto too little memory, because that is where output visibly degrades. Which model, though? That is a separate decision with its own moving parts, and the open-weight landscape shifts monthly. Rather than chase version numbers, start with a well-known family and a size your hardware can hold: Meta's Llama, Alibaba's Qwen (including its Qwen Coder variants), Google's Gemma, Mistral, or DeepSeek. For a current shortlist by use case, see our guide to the [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms), and note that several of the strongest options are [open-weight families like Qwen and DeepSeek](https://geotoolbox.ai/blog/chinese-ai-models-compared) you can run entirely offline. The rule that ages well: pick a family, pick a size that fits, download the Q4_K_M file. ## Run Your First Model in 15 Minutes With Ollama, the whole process is short, and most of the fifteen minutes is the model downloading in the background. 1. **Install Ollama.** Download it from [ollama.com](https://ollama.com/download) and run the installer. It sets up a background service and adds the `ollama` command to your terminal. 2. **Pull and run a model.** In a terminal, type `ollama run llama3.2`. Ollama downloads the model on the first run, then drops you into an interactive chat right there in the terminal. Swap `llama3.2` for any model in its [library](https://ollama.com/library); a 3B model is a safe first pull on modest hardware. 3. **Chat.** Type a message and press enter. That is a large language model answering from your own machine, no account and no internet required.
![A local LLM data flow: your prompt goes to a runner, then the model in your computer's memory, and the answer returns, with nothing sent to the cloud.](/blog/run-llm-locally/local-llm-data-flow.png)
With a local model, the whole loop stays on your machine. Nothing is sent to a cloud provider.
Prefer a familiar window over a terminal? Point [Open WebUI](https://docs.openwebui.com/) at Ollama and you get a browser-based, ChatGPT-style interface with chat history and document upload for asking questions over your own files, all still running locally. LM Studio users get the same thing built in. First-run downloads are large. A small model is a few gigabytes, a 30B model can be twenty or more, and a 70B file alone can top 40 GB. Keep several gigabytes free for a small model and considerably more for the bigger tiers, and use an SSD, not a spinning drive. Run `ollama list` to see what you have pulled and `ollama rm ` to clear the ones you are done with. ## Use Your Local Model for Coding This is where a local model earns its keep, and it is the part most guides skip. The key is that Ollama and LM Studio expose an **OpenAI-compatible API** on your machine. Anything that can talk to OpenAI can talk to your local model instead by pointing at a different address. Ollama serves that endpoint at `http://localhost:11434/v1`. Because the request never leaves your computer, the API key is not checked, so any placeholder value works. For **Cursor** or the **Continue** extension in VS Code and JetBrains, the setup is the same idea in both: add a model provider, set the base URL to your local endpoint, and choose a model you have pulled (a Qwen Coder model is the usual pick for this). [Continue](https://www.continue.dev/) has native Ollama support in its config, so it is often the smoothest path to a local coding assistant. For a wider set of options here, we keep a running list of [open-source tools for Claude Code](https://geotoolbox.ai/blog/claude-code-open-source-tools). **Claude Code has one gotcha worth knowing before you waste an hour on it.** Claude Code speaks Anthropic's message format, not OpenAI's, so it does not use the same `/v1` endpoint as Cursor or Continue. Ollama exposes a separate Anthropic-compatible endpoint for exactly this, and the simplest path is to let it connect everything for you with `ollama launch claude`. If you set it up by hand instead, export these variables and add them to your shell profile so they persist: `ANTHROPIC_BASE_URL` pointed at your local Ollama, `ANTHROPIC_AUTH_TOKEN` set to any value, and `ANTHROPIC_API_KEY` set to an empty string. Do not skip that last one: a cloud Anthropic key already sitting in your environment can otherwise override the local setup, which is the exact dead end this section is about. See [Ollama's documented Claude Code integration](https://docs.ollama.com/integrations/claude-code) for the full walkthrough. Be honest with yourself about the result. A local 7B to 32B coding model is not GPT-5 or Claude Opus. It handles autocomplete, boilerplate, and private codebases you cannot paste into a cloud tool, and it costs nothing per token after the hardware. Treat it as a strong second seat for privacy-sensitive or offline work. ## Why Is It Slow? Fixing Common Problems The most common complaint is speed, and the cause is almost never a weak GPU. It is silent offload. When a model does not fully fit in VRAM, the runner quietly moves part of it into system RAM, which is far slower, so generation crawls. You do not need new hardware. Pick a model or quantization that fits. Diagnose it in one command. Run `ollama ps` while a model is loaded and look at the processor column: if it says anything other than 100% GPU, part of the model is on the CPU. From there: - **Pick a smaller model or quant** so the whole thing fits in memory. - **Lower the context length.** A longer context holds more of the conversation but eats memory and slows generation; along with [settings like temperature](https://geotoolbox.ai/blog/ai-temperature), it is one of the main levers you control. - **Keep the model resident** so it is not reloaded on every request. A realistic baseline: interactive chat feels fine above about 15 tokens per second and sluggish below 5. If a model that should fit your memory, say 7B to 32B, is crawling in single digits on a capable GPU, you are almost certainly offloading. The largest 70B-class models are slower by design and are the exception. The other classic failure is **the GPU not being used at all on Windows**, especially inside WSL2. The usual mistake is installing NVIDIA drivers *inside* WSL2, which breaks the passthrough. The host Windows driver already provides GPU access to WSL2, so do not install a second one there. AMD support through ROCm is thinner than NVIDIA's; if you are on AMD or fighting WSL, a Mac with Apple Silicon is the least painful path by a wide margin. Local runners also default to a fairly small context window, which is why a local model can seem to "forget" earlier in a conversation than ChatGPT does. You can raise it, with Ollama's `num_ctx` setting or by choosing a longer-context model, but a bigger context uses more memory and slows generation. It is a trade-off, not a free upgrade. That trade-off has a sharp edge worth watching: the **context cliff.** A model that fits comfortably for a short chat can run out of memory when you paste a long document, because a bigger context grows the memory it needs. This is also why the 32B and 70B ceilings in the hardware table assume short-to-moderate context; a very long prompt eats into the headroom those models need. If a model crashes only on long inputs, that is why. Trim the input or raise your memory headroom. ## Is Running an LLM Locally Worth It? For privacy, offline access, bulk or repetitive work, and learning, yes. As a cheaper drop-in for a frontier model on your toughest tasks, no. The cost case is real but conditional. If you already run a capable GPU or a high-memory Mac, local inference is effectively free after electricity, and for heavy or bulk use it undercuts a per-token API bill quickly. If you would have to buy hardware specifically for this, a one-time GPU in the few-hundred-dollar range competes with tens of dollars a month in [a ChatGPT subscription or the API](https://geotoolbox.ai/blog/chatgpt-pricing) only once your usage is high enough. For a light, occasional user, paying for the cloud is cheaper and less hassle. Buy hardware for the privacy and control, not to save a few dollars. Where local wins outright is the work you cannot or should not send to a vendor: confidential client files, proprietary code, anything under a policy that forbids pasting into cloud tools, and any task you need to run with no internet at all. It is also the best way to build real intuition for how these systems behave. That intuition pays off directly if you work in marketing or SEO. The open-weight models you can now run on a laptop are increasingly the same ones answering questions about brands inside AI search, so running one locally is a cheap way to watch how they read and summarize content. ## Frequently Asked Questions ### How much RAM do you need to run an LLM locally? Around 8 GB runs 7B to 8B models at Q4_K_M quantization, 16 GB is the comfortable everyday range and handles 13B to 14B, and 24 GB or more opens up models in the 32B class. On a discrete GPU the relevant number is VRAM; on an Apple Silicon Mac it is unified memory. A rough guide is about half a gigabyte for the weights of each billion parameters, plus headroom for the runtime and your conversation. ### Do I need a GPU to run an LLM locally? No, but it helps a lot. A modern CPU with enough RAM can run small models at roughly 3 to 8 tokens per second, which is fine for background tasks and slow for live chat. A GPU, or an Apple Silicon Mac's unified memory, is what makes local models feel responsive. ### Is it worth running LLMs locally? It is worth it for privacy, offline use, bulk or repetitive jobs, and learning. It is not worth it as a cheaper replacement for a frontier model on complex reasoning, where a cloud model still wins. If you already own capable hardware, it costs nothing per token; if you would buy hardware just to save money, the math only works at high usage. ### What laptop can run an LLM locally? Any laptop with 16 GB of memory can run useful 7B to 8B models. Apple Silicon MacBooks (M-series) are especially strong because their unified memory lets the GPU use a large shared pool, so a high-memory Mac punches well above a PC laptop with the same total memory. ### Is running an LLM locally actually private? Yes, in the sense that prompts are processed on your machine and are not sent to a model vendor. A couple of caveats: some apps still make network calls for updates or optional web search, and if you expose the local server to your network, other devices could reach it. Keep the server bound to your own machine and disable optional cloud features if privacy is the point. ### Which local model should I start with? Start with a well-known family at a size your memory can hold rather than chasing the newest version number: a small Llama or Qwen model at Q4_K_M is a safe first pull, and a Qwen Coder model is a good default for coding. Once one is running, try a larger size from the same family to feel the quality and speed trade-off. ## Where to Go From Here The fastest way to understand local models is to run one. Install Ollama or LM Studio, pull a small model that fits your memory, and give it a real task tonight. Fifteen minutes later you will know more about what these systems can and cannot do than any benchmark chart will tell you. Once you have seen how a model reads and summarizes text, the natural next question for anyone doing marketing or SEO is whether the AI systems built on these models can actually find and cite your own site. That is the problem geotoolbox is built for. Start by checking whether AI crawlers can reach your pages at all with our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker). ## Sources - Ollama Quickstart and Docs - Ollama, 2026 - `docs.ollama.com/quickstart` - Running local models with Claude Code - Ollama, 2026 - `docs.ollama.com/integrations/claude-code` - Hugging Face model hub - Hugging Face - `huggingface.co` - LM Studio - Element Labs - `lmstudio.ai` - Continue (local AI code assistant) - Continue - `continue.dev` - llama.cpp inference engine - ggml-org, GitHub - `github.com/ggml-org/llama.cpp` --- ## Semrush vs Ahrefs (August 2026): What the Data Says > We ran Semrush and Ahrefs on the same three sites, twice, twelve days apart, and checked both against Google Search Console. Here's where they disagree, and which numbers to trust. - Canonical: https://geotoolbox.ai/blog/semrush-vs-ahrefs - Published: 2026-07-21 · Updated: 2026-08-22 Most Semrush vs Ahrefs comparisons hand you a feature checklist and a winner. This one starts with a problem you have probably seen yourself: point both tools at the same website on the same day and the numbers don't match. Not by a little. On one of our sites, Ahrefs reports 42 organic keywords and Semrush reports about 1,500. So we ran the test on real sites. Both tools, on three sites we manage, first in July 2026 and again on August 1, cross-checked against the one source that isn't an estimate: Google Search Console. Running it twice tells you how stable each tool's numbers are; the Search Console check tells you which of them ran closer to reality. We are a Semrush affiliate, so we have an incentive to be liked here. That is exactly why this piece leans on our own account data and Google Search Console rather than adjectives, shows where Ahrefs wins, and flags where Semrush's numbers are the less accurate of the two. The three-site keyword and traffic figures, plus the geotoolbox.ai backlink and AI-citation figures, come from our own accounts; pricing and third-party test figures are cited to their sources below, so you can run the same test and check us. Buying Semrush through our links earns us a commission at no extra cost to you. ## Semrush vs Ahrefs: The Short Answer **Semrush is the all-in-one marketing platform of the two — SEO, AI visibility, paid search, content, local, and social in one account — and the better default for most buyers. Ahrefs is the focused specialist's tool, built around backlinks and organic research.** If your work spans channels, or you run an agency, a marketing team, or a young site you want to watch grow, pick Semrush; it also reported far more ranking keywords on every site we tested, though broader coverage is not better accuracy. Pick Ahrefs if you are the SEO specialist in the room, the person whose whole job is links, rankings, and site audits, and you value a cleaner interface and more conservative numbers over breadth. If Semrush is where you land, the [7-day free trial](/go/semrush-seo?ref=semrush-vs-ahrefs-answer) is enough time to run every test in this article on your own site. A couple of things hold no matter which you choose. First, the tools will disagree, often dramatically, because they track different keyword databases, SERPs, and traffic models. Second, neither number is the truth. For your own Google traffic, [Google Search Console](https://search.google.com/search-console/about) is the closest thing to ground truth, and both tools are best treated as competitive-intelligence estimates, not measurement. The rest of this article shows how far apart they run on our own sites. ## We Ran Both Tools on the Same Three Sites We pulled fresh numbers from our own paid Ahrefs and Semrush accounts for three sites we manage: geotoolbox.ai (our own small, new AI-visibility tool) plus two client sites we have kept anonymous — an established, mid-size manufacturer and a large marketplace. Same domains, same day, first in mid-July 2026 and again on August 1. Here is what each tool said about organic keywords on the August pull. **How we pulled this, so you can weigh it.** Three sites is an illustration, not a study. Ahrefs figures are all-locations (our account is on the Lite tier, but plan tier limits row exports, not the headline counts, which come from the dashboard overview). Semrush's keyword figures use its worldwide database for geotoolbox and its US database for the two US-focused sites (the page counts below are Ahrefs all-locations against Semrush's US database). That scope is not perfectly matched, so treat the exact multipliers below as rough direction. The direction is what holds up.
![Scorecard comparing Ahrefs and Semrush for geotoolbox.ai on the same day in August 2026: Ahrefs shows 42 organic keywords, 375 estimated traffic, 121 referring domains, and DR 1.9; Semrush shows about 1,500 organic keywords, 1,700 estimated traffic, 44 referring domains, and Authority Score 11.](/blog/semrush-vs-ahrefs/geotoolbox-scorecard.png)
geotoolbox.ai in both tools on the same day. Ahrefs: 42 keywords. Semrush: about 1,500. Neither is lying; they index the web differently.
SiteSizeAhrefs organic keywordsSemrush organic keywordsSemrush ÷ Ahrefs
geotoolbox.aiSmall / new42~1,500~36x
Client A — manufacturerMid-size374~2,000~5.3x
Client B — marketplaceLarge26,60076,000~2.9x
One thing holds across all three: Semrush reports more keywords on every site, even on the two where it was restricted to a single US database. The size of the gap is where we are more cautious. In our sample it was widest on the smallest site (about 36 times) and narrower on the large one (about three times), which fits the idea that Semrush's database catalogs more of the long-tail and low-volume terms a young site picks up first. But that gradient rests on three points, and Semrush ran worldwide on the smallest site while restricted to the US on the two larger ones — a scope difference that happens to track site size — so test it on your own accounts before you rely on it. The same pattern shows up in pages, with a caveat worth stating plainly because geotoolbox.ai publishes in both English and French. Ahrefs's all-locations report surfaces **22 ranking pages**; Semrush's US database shows **106** (its US index does not see the site's French pages at all); and Google Search Console had **around 225 URLs indexed** in July — a count spanning every English and French blog post, glossary entry, and feature page the site publishes, not 225 articles. These are not the same measurement — ranking pages in one geographic scope versus Google's full multilingual index — so read them as three differently-framed slices of one site, not a clean ratio. What they agree on is the part that matters: no third-party tool sees the whole footprint, and Search Console comes closest to the full picture.
![Ahrefs Top Pages report for geotoolbox.ai showing 22 ranking pages and 373 total organic traffic across those pages in August 2026.](/blog/semrush-vs-ahrefs/ahrefs-pages.png) ![Semrush Top Pages summary for geotoolbox.ai in its US database: Organic Pages 106 (up 55.88%), with the pages count climbing through mid-2026.](/blog/semrush-vs-ahrefs/semrush-pages.png)
Ranking pages for geotoolbox.ai: Ahrefs 22 (all-locations, top; its 373 traffic figure sums those pages, against 375 site-wide), Semrush 106 (US index only — it excludes the site's French pages, and its "Organic Traffic 7" tile is the US-only slice of the worldwide 1,700 estimate used elsewhere). Google Search Console had around 225 URLs indexed in July, counting both languages' posts, glossary entries, and feature pages.
The reason is structural: each tool tracks a different universe of keywords, countries, and SERPs, on different refresh schedules and with different thresholds for what counts as a ranking keyword. A keyword count is database coverage, not verified visibility, accurate live positions, or traffic. Which raises the obvious question. ## The Same Test, Twelve Days Later: How Each Tool's Numbers Move Over twelve days, Semrush added roughly 400 keywords (+36%) and 14 ranking pages for our test site while Ahrefs added seven keywords (+20%) and two pages — and Ahrefs sharply revised its traffic estimate downward. A single same-day snapshot hides that kind of movement, so we kept both accounts pointed at geotoolbox.ai and re-pulled everything on August 1. (The client sites moved too, but a young site is where the difference between the two tools is most visible, so that is the one under the microscope.)
geotoolbox.aiAhrefs (Jul 20)Ahrefs (Aug 1)Semrush (Jul 20)Semrush (Aug 1)
Organic keywords3542~1,100~1,500
Ranking pages202292106
Authority scoreDR 1.7DR 1.9AS 7AS 11
Referring domains641212744
Est. organic traffic~1,500375~1,0001,700
Both tools grew their picture of the site over the window; what differs is resolution. Semrush's catalog registered the site's recent publishing as hundreds of new keyword entries and a visible Authority Score move from 7 to 11, while Ahrefs's stricter thresholds recorded the same period as a handful of keywords. (DR and AS are separate proprietary scales that don't convert, so read each tool's movement only against itself.) If you publish weekly and want the dashboard to reflect that work while it is happening, Semrush is the one that shows it. Ahrefs has two answers of its own, though. Its backlink crawler was the faster mover, going from 64 to 121 referring domains for the site against Semrush's 27 to 44, exactly what you would expect from the stronger link index. And its traffic estimate dropped from about 1,500 to 375, which looks erratic in isolation; hold that thought, because the next section, where both tools meet Search Console, is where it resolves. On this site, over this window, the trade reads: **Semrush shows you the progress first; Ahrefs waits for proof, then corrects hard.** One site is not a pattern, but it is the behavior to watch for. For a new site owner deciding whether the content plan is working, the tool that registers movement is the more useful weekly companion, as long as you keep treating the traffic number as a model, not a meter. ## Coverage Isn't Accuracy: What Google Search Console Showed The extra keywords Semrush reports could be real reach or pure noise. To find out which, we pulled the actual clicks Google Search Console recorded for each site over the same 30-day window. GSC counts the real Google organic clicks your site received, not visits from every channel, which makes it the best first-party benchmark for exactly the thing the tools are trying to estimate.
SiteAhrefs estimated trafficSemrush estimated trafficGSC actual clicks (30 days)
geotoolbox.ai375~1,700382
Client A — manufacturer2,7005,1003,703
Client B — marketplace316,000397,30083,499
![Bar chart comparing Ahrefs and Semrush estimated monthly organic traffic against actual Google Search Console clicks for three sites in August 2026: both tools overshoot the largest site; on the mid-size manufacturer Ahrefs undershoots while Semrush overshoots; on the smallest site Ahrefs's 375 estimate lands almost exactly on the 382 real clicks while Semrush reads 1.7K.](/blog/semrush-vs-ahrefs/estimate-vs-gsc.png)
Estimated organic traffic vs the clicks Google actually reported, August 2026. Ahrefs ran closer to Search Console on all three sites, though both suites overshoot the marketplace several times over.
The gaps are large. Note that these tool estimates are modeled monthly organic traffic while GSC is actual clicks over 30 days, so the two are not a perfect like-for-like; the point is the size of the gap, not a precise error rate. On the large site both tools ran 3.8 and 4.8 times over Google's 83,499 clicks: Ahrefs estimated 316,000 and Semrush 397,300. The mid-size manufacturer split them, with Ahrefs undershooting (2,700 vs 3,703 real) while Semrush ran high at 5,100. And the smallest site is where the August re-pull got interesting: Ahrefs, which had estimated around 1,500 in July, revised down to 375 — within a handful of clicks of the real 382 — while Semrush's 1,700 estimate sat about 4.5 times above reality. Semrush read higher than Ahrefs on all three sites, and Ahrefs ran closer to Search Console on all three — dramatically so on the smallest — at the cost of that big mid-July swing. None of that makes either tool bad. It makes their "organic traffic" a modeled estimate, not a click counter. It also matches the most-cited public test on the question: a 2022 study in which Ahrefs checked its own US organic traffic estimates against Google Search Console on [1,635 websites](https://ahrefs.com/blog/traffic-estimations-accuracy/) and found a median deviation of 49.52% for Ahrefs and 68.36% for Semrush. One caveat, since we are the Semrush affiliate: Ahrefs ran and published that study, so read it as an interested party's finding, not a neutral referee. Its direction still lines up with what we saw — Semrush read higher on all three of our sites, and in that 2022 sample its estimates deviated more — and with what practitioners repeat constantly: the estimates are useful for spotting trends and comparing competitors, never for reporting your own traffic. For that, open Search Console. The gap has a narrow practical use. Your own site's ratio of real Search Console clicks to a tool's estimate can roughly sanity-check that tool's number for a competitor, but only a competitor of similar size and market. Emphasis on rough: our three test sites show the error swinging from a near-exact match to nearly a five-times overshoot, so a ratio borrowed from one site will not transfer cleanly to a site of a different size or age. It is a gut check at best. ## Keyword Research: Database Breadth vs Click Realism The personality split shows up hardest here. Semrush's Keyword Magic Tool is built for breadth. It returns huge lists grouped into topic clusters with search intent attached, plus a Personal Keyword Difficulty score tuned to your domain. If your job is to find every angle on a topic and map coverage, Semrush surfaces more to work with, the breadth behind the larger keyword counts we saw.
![Semrush Keyword Magic Tool for the seed 'ai visibility tools', showing 4,824 keywords grouped into topic clusters with search intent, keyword difficulty, and volume columns.](/blog/semrush-vs-ahrefs/semrush-keyword-magic.png)
One seed keyword in Semrush's Keyword Magic Tool fans out to 4,824 terms, grouped by cluster, intent, and difficulty.
Ahrefs Keywords Explorer is built for realism. Its signature metric is Traffic Potential: Ahrefs's estimate of the monthly organic traffic the current top-ranking page for a term gets from every keyword it ranks for, not just the head keyword's volume. It also reports a Clicks metric, its estimate of the average monthly search-result clicks a keyword generates, which accounts for zero-click searches as well as searches that produce several clicks. On a query where an AI Overview or featured snippet eats most of the clicks, that distinction is the difference between chasing a keyword and skipping it. The database-size debate is mostly noise. Comparison articles quote keyword-database figures for the same tool that differ severalfold depending on who is writing, which tells you these numbers are unreliable marketing counts, not audited facts. Both tools index tens of billions of keywords; that is plenty. What differs is philosophy: Semrush is usually described as the more expansive, higher-volume counter and Ahrefs as the lower, more conservative one, which is the shape our own three-site test showed. Both approaches are defensible, and the AI answers you see today increasingly repeat both sets of figures as if they were settled, which they are not. Verdict: Semrush for breadth and content planning, Ahrefs for judging whether a keyword will actually send clicks. ## Backlinks: Where Ahrefs Still Earns Its Reputation Ahrefs built its name on backlinks, and it is still the default for link work, but the "bigger index" story has flipped on paper. By each vendor's own self-reported (and take-with-salt) count, Semrush now advertises the larger raw number — around 43 trillion links, against the external-backlinks history Ahrefs publishes on its own big-data page — though the two are counted under different definitions, so treat both as marketing figures. In a third-party ten-domain test, Style Factory found Semrush reported more referring domains on nine of them, including Amazon (4.4 million vs Ahrefs's 2.2 million). Then look at our own smallest site and it flips the other way. For geotoolbox.ai, Ahrefs finds 121 referring domains and Semrush finds 44, and over our two-pull window Ahrefs picked up the site's new links faster (the movement table above has the exact deltas). Same site, and the tool with the "smaller" index found nearly three times as many referring domains. The counts disagree in both directions, and raw index size tells you very little. Practitioners often reach for a neutral yardstick here: [Cloudflare Radar's bot directory](https://radar.cloudflare.com/bots/directory), which ranks crawlers by the traffic Cloudflare actually observes across its network. When we checked on August 7, 2026, AhrefsBot was the highest-placed SEO crawler on the list at number 13 overall by traffic volume, one place ahead of Semrush's backlink crawler, SemrushBotBacklinks, at 14. It is a snapshot, and the order moves. What separates them is quality and history. Ahrefs refreshes its link index constantly and keeps years of historical data, and its crawler has long been the practitioner favorite for link audits, disavow work, and finding a competitor's best links. A huge index full of dead or spammy links is worse than a smaller, cleaner one. If backlinks are the core of your work, Ahrefs remains the safer default, and our [best link building platforms](https://geotoolbox.ai/blog/best-link-building-platforms) guide covers the outreach tools that sit on top of whichever index you use. If you want backlinks as one part of a wider suite, Semrush's data is now genuinely competitive on volume. ## Everything Else: Audits, Rank Tracking, PPC, Local, Content Outside keywords and links, the split is simpler: Semrush does more things, Ahrefs does fewer things with a cleaner interface. This is the all-in-one platform versus focused specialist trade-off, and it decides most purchases.
JobAhrefsSemrushEdge
Site auditClean, solid technical crawlBroader issue coverage and guided fixes, in our useSemrush
Rank trackingGoogle, desktop and mobile; weekly by default on lower tiersDaily updates; Google, Bing, and moreSemrush
Paid search / PPCBasic ad dataFull PPC, ad history, keyword gapSemrush
Local SEONot a focusListings, review, map-pack tools (add-on)Semrush
Content toolsBasic AI writing helpersSEO Writing Assistant, content templatesSemrush
Ease of useCleaner, faster to learnPowerful but a steeper curveAhrefs
The table reads as a Semrush sweep, and on feature count it is one. But breadth has a cost: Semrush's local, agency, and AI features are often paid add-ons on top of the base plan, and the interface carries the weight of doing everything. Ahrefs does less on purpose, and people who live in it every day tend to prefer that focus. If your work is genuinely just organic search and links, you are paying for modules you will never open. One ownership change is worth knowing before you commit. Semrush is now an Adobe company: [Adobe completed its acquisition on April 28, 2026](https://news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition), and says it will fold Semrush's data into its enterprise marketing stack. If your team already runs on Adobe, that direction is a genuine plus for the all-in-one case. Some practitioners in SEO forums went the other way after the announcement, renewing with Ahrefs — now the only independent of the two — on pricing-and-roadmap wariness under new ownership. Neither reaction changes what the tools measure today, but it is part of the 2026 buying picture. ## AI Search Visibility: Where Semrush Pulled Ahead in 2026 This is a big reason Semrush now wins the "which one platform" question on integration and entry price for anyone whose audience is moving into AI answers. With a growing share of searches ending inside an AI answer instead of a blue link, whether you show up when ChatGPT, Gemini, or Google's AI Mode answers a question in your market is no longer a bonus metric. Both tools track it. They do not do it the same way. Semrush's answer is the [AI Visibility Toolkit](/go/semrush-ai?ref=semrush-vs-ahrefs-ai), and it is the better-packaged, better-integrated of the two. It scores your brand with an AI Visibility Score (a 0–100 benchmark of how often AI answers mention you versus the median of your top competitors) and a Share of Voice percentage, across ChatGPT, Google AI Mode, Google AI Overviews, Gemini, and Perplexity, per Semrush's plan documentation. Prompt Research surfaces the prompts worth owning and who competes on them; a cited-sources breakdown shows which pages the engines actually pull from; daily Prompt Tracking follows 25 custom prompts so you can benchmark competitors and watch shifts over time; and since Semrush's 2026 repositioning the toolkit has grown into a full AI Visibility section of the app, with brand-performance, perception, and [brand-sentiment](https://geotoolbox.ai/blog/ai-brand-sentiment) reports alongside the tracking. If you already run a Semrush SEO plan, that AI layer is a separate purchase (pricing below); the [Semrush One bundles](/go/semrush-one?ref=semrush-vs-ahrefs-ai) from $199/mo fold it in. Either way it sits in the same account as your keyword, backlink, and content data, which is the unified AI-plus-SEO pitch.
![Semrush AI Visibility Main Metrics for geotoolbox.ai: Mentions 0, Citations 39 (up 550%), and Cited Pages 25 (up 400%), charted rising from June to July 2026.](/blog/semrush-vs-ahrefs/semrush-ai-toolkit.png)
Semrush's AI Visibility Toolkit on our own geotoolbox.ai account: citations (39) and cited pages (25) tracked over time. The percentage badges use the dashboard's own shorter lookback window, not the July-to-August comparison discussed in the text.
Ahrefs answers with Brand Radar, and on raw engine coverage it is actually the wider of the two: it adds Microsoft Copilot, Grok, and Claude on top of the same ChatGPT / AI Overviews / AI Mode / Gemini / Perplexity set (plus non-AI discovery surfaces like TikTok), backed by a 472M-plus prompt database, with its own competitor benchmarking and custom prompt tracking. That width comes with two asterisks: the Grok index is currently paused (Ahrefs says it cannot gather new Grok data right now), and Claude queries burn extra checks. Where Brand Radar loses is packaging and price.
AI-visibility featureSemrush AI Visibility ToolkitAhrefs Brand Radar
Engines trackedChatGPT, Google AI Mode + AI Overviews, Gemini, PerplexitySame set plus Copilot, Grok (index currently paused), and Claude (and discovery surfaces like TikTok)
Brand scoreAI Visibility Score (0–100) + Share of VoiceMentions and share metrics
Prompt trackingDaily, 25 custom prompts on the base AI-toolkit tierDaily or monthly, quota depends on package
Competitor benchmarkingYes, plus brand sentimentYes
Cited-source breakdownYesYes
Cost to actually use it$99/mo per domain (same price monthly or annually), or bundled with SEO in Semrush One from $199/moSmall quota inside Lite ($129/mo); prompt-only add-ons from $50/mo; full per-platform indexes at $199/mo each, or $699/mo for all platforms with 2,500 prompt checks
Two caveats before you read that table as a clean Semrush sweep. First, the AI Visibility Toolkit is sold separately — it is not folded into Pro, Guru, or Business, so it is a real line on the bill; where it wins is packaging and the entry price for the full visibility suite, not broader engine coverage. Note the units differ, too: Semrush's $99 is per domain and multiplies across a client roster, while Ahrefs's $199/$699 is per platform index, so a multi-brand agency should run its own math. Second, both tools model AI visibility with simulated prompts rather than watching real user sessions, so treat their scores the way you treat their traffic estimates: excellent for your own trend line and for comparing competitors, not a literal count of every time a human saw your name. (One small asterisk from our own account: the toolkit's Distribution by LLM panel showed data for four of the five documented engines — Perplexity had nothing for our domain yet.) We also ran this layer on a site we own. Point both tools at the same brand and they surface AI citations a Google-clicks report never registers. geotoolbox.ai earns around 380 Google clicks a month after roughly two and a half months of publishing — still a tiny number — yet both tools flag it turning up in AI answers. The counts, as of August 1: Ahrefs logged **95 ChatGPT responses across 15 of its pages**, up from 29 across 9 in July. Semrush's AI Visibility Toolkit logged **39 citations across 25 cited pages**, up from 31 across 19. Ahrefs's Site Explorer overview also now carries a free per-platform AI index, a genuinely useful addition that showed 108 Copilot and 64 Perplexity responses for the same site. Those are different units and don't compare head-to-head, but together they make the point concrete: for an AI-first brand, keyword count is a progress signal, while whether engines actually cite you is a separate question that AI-visibility tools are built to measure. Then we pulled the same site's first-party data from Bing Webmaster Tools, whose AI Performance report logs citations made inside Microsoft's Copilot ecosystem. It counted **15,300 citations across an average of 24 cited pages in three months** — orders of magnitude above the 95 responses and 39 citations the two suites logged. Bing calls its own figure a sample, and it covers a different engine, period, and unit than the suites' simulated prompts, so this is not a scoreboard. What it shows is scope: the suites sample a fixed prompt set across their engine panels (Ahrefs's includes Copilot, but samples it), while Bing logs live citations inside its own ecosystem. No single source sees everything.
![Bing Webmaster Tools AI Performance report for geotoolbox.ai showing 15.3K total citations and 24 average cited pages over three months, sourced from Microsoft Copilots and partners.](/blog/semrush-vs-ahrefs/bing-ai-performance.png)
The same site in Bing Webmaster Tools: 15,300 citations logged in Microsoft's Copilot ecosystem over three months — a first-party sample, measured differently than the suites' simulated prompts.
So the workflow is the one this article keeps arriving at: lean on the suites (Semrush's toolkit or Ahrefs's) to track your trend and watch competitors, and cross-check against the first-party data the platforms report — Bing Webmaster Tools now, Google's as it lands — plus an independent cross-engine read, remembering each is its own partial slice. To go deeper on the AI side specifically, we cover how to read [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice), why engines fan one prompt into [dozens of sub-questions](https://geotoolbox.ai/blog/query-fan-out) (exactly why wide keyword coverage still helps here), and how the dedicated [AI visibility tools](https://geotoolbox.ai/blog/best-ai-visibility-tools) stack up. ## Pricing: What You'll Actually Pay in 2026 Sticker prices are close at every tier. The real difference is what sits on top: Ahrefs's cheaper plans meter usage with credits, while Semrush charges add-ons for extra users, local, and AI visibility. Here are the core plans, verified against each vendor's pricing page in August 2026. Confirm current numbers before you buy, since both change often.
TierAhrefsSemrush (SEO)
Cheapest full planLite, $129/mo (limited Starter from $29)Pro (now sold as "SEO"), $139/mo
MidStandard, $249/moGuru, $249.95/mo (off public pricing, see below)
HighAdvanced, $449/moBusiness, $499.95/mo (off public pricing, see below)
Free / trialAhrefs Free, formerly Ahrefs Webmaster Tools (own verified sites); no time-boxed trial7-day free trial + limited free account
A few things the table hides. Semrush's public pricing page — now titled "SEO & AI Search Plans and Pricing" — shows only an entry "SEO" plan (the classic Pro renamed, and repriced from $139.95 to $139) next to its One bundles; Guru and Business no longer appear on public pricing and persist in the docs and for existing subscribers. Our [full Semrush review](https://geotoolbox.ai/blog/semrush-review) breaks down what that changes. Ahrefs's entry Lite plan is affordable but credit-metered (1,000 credits per user), and those credits burn fast in a deep research session, while its higher tiers move to unlimited credits but keep other export and project limits. Semrush's headline price is the base; extra seats are a paid add-on, and its AI Visibility Toolkit is a separate line item at $99/mo per domain, the same price on either billing term, or bundled with SEO in the Semrush One tiers ($199, $299, or $549/mo), so an agency's real bill climbs above the sticker. Both suites discount annual billing (Semrush by up to about 17%). And Ahrefs offers no classic time-boxed free trial (its free access is Ahrefs Free, the renamed Webmaster Tools, for sites you own), which matters if you want to test before committing. If Semrush's breadth is the fit and you want to run this same three-site test on your own accounts, its [SEO Toolkit](/go/semrush-seo?ref=semrush-vs-ahrefs-pricing) starts at $139 a month (less billed annually) and there is a [7-day free trial](/go/semrush-seo?ref=semrush-vs-ahrefs-trial) to try it first. On a tight budget, start with our [best free SEO tools](https://geotoolbox.ai/blog/best-free-seo-tools) guide before paying for either. ## So Which One Should You Buy? Match the tool to how you work rather than to a scoreboard. If Semrush is the direction you are leaning, our [full Semrush review](https://geotoolbox.ai/blog/semrush-review) goes deeper on its 2026 pricing, the AI Visibility Toolkit, and the billing complaints worth knowing before you subscribe.
If you are a...PickWhy
New or small site ownerStart free, then SemrushBegin with Search Console and Keyword Planner (free). When you pay, Semrush shows far more of your early presence, useful for direction, though that bigger number is what its database sees, not visits you are getting. Ahrefs has a limited $29 Starter plan, but for real research both suites' full plans cost about the same.
SEO specialist or link builderAhrefsIf SEO is the whole job — link audits, disavows, organic research, technical crawls — the focused tool wins: cleaner interface, trusted backlink data, conservative numbers, nothing you won't open.
Agency, marketing team, or generalistSemrushSEO, PPC, content, local, AI visibility, and reporting in one vendor ecosystem (some as paid add-ons) beats stitching separate tools together. If you'd rather hire the work out, start with our best AI SEO agencies guide.
AI-first / GEO-focused siteSemrush + its AI Visibility ToolkitOne vendor covers the search layer and the AI-answer layer: Visibility Score, Share of Voice, and daily prompt tracking beside the SEO data.
The plain verdict: for most people buying one tool in 2026, Semrush is the default, because most of them are not pure SEOs. It is the all-in-one marketing platform of the pair, it shows a small site more of its own footprint sooner — remembering that broader coverage is not better accuracy — and its AI-visibility layer comes from the same vendor as everything else. Ahrefs is the better buy for the SEO specialist: if you spend your days in backlink profiles and organic research and want conservative numbers in an uncluttered tool, its focus is exactly the point. Neither will report your real traffic, so keep Search Console open next to whichever you choose. One last thing, because it is the reason we ran this test at all. If your visibility is moving into AI answers, the tool question is only half of it. Of the two, Semrush has pushed AI visibility furthest into its platform: its AI Visibility Toolkit pulls citation tracking, Share of Voice, and daily prompt monitoring into the same account as the SEO data. For a site like ours, whose Google clicks are tiny but whose name turns up in AI answers, that layer is the one worth watching most. Run the [7-day Semrush trial](/go/semrush-seo?ref=semrush-vs-ahrefs-closing) to pull your own numbers, then check them against Search Console and an independent AI-citation read — the disagreement between the three tells you more than any single tool will. ## Frequently Asked Questions ### Which is better, Semrush or Ahrefs? For most people buying one tool, Semrush is the better pick because it is an all-in-one marketing platform: SEO, paid search, content, local, and AI visibility in one account, and it surfaces more keywords on every site we tested. Ahrefs is better for SEO specialists whose work is led by backlinks and organic research and who prefer a cleaner interface with more conservative numbers. Neither reports your real traffic, so pair whichever you choose with Google Search Console. ### Can you trust Semrush's traffic numbers? Trust them for relative comparison and trend direction, not as literal visits. In our August 2026 test across three sites, Semrush's organic-traffic estimate ran higher than Ahrefs's on all three and overshot actual Google Search Console clicks on each; Ahrefs overshot the largest site badly too, but landed within a few clicks of reality on our smallest site after a sharp downward revision. Both tools model traffic from rankings and click curves; only Search Console counts your real clicks. ### Why do Semrush and Ahrefs show different keyword counts for the same site? Because they track different keyword databases, update schedules, and definitions of a ranking keyword. Semrush cataloged more terms on every site we tested; we did not test which query types drove the difference. In our data the gap ranged from about 3 times more on a large site to 36 times more on a small new site. On our smallest site the gap widened between pulls (31x to 36x, with Semrush adding roughly 400 keywords to Ahrefs's seven); on the two larger sites it narrowed slightly. ### Does Semrush track AI search visibility? Yes. Semrush's AI Visibility Toolkit tracks how your brand appears across ChatGPT, Google AI Mode and AI Overviews, Gemini, and Perplexity, with an AI Visibility Score, Share of Voice, Prompt Research, cited-source breakdowns, and daily tracking of 25 custom prompts for competitor benchmarking. It is priced separately ($99/mo per domain, the same price monthly or annually, or bundled with SEO in Semrush One from $199/mo), not part of the standard SEO plans. Ahrefs offers similar tracking through Brand Radar across an even wider engine set (it adds Copilot, Grok, and Claude, plus non-AI discovery surfaces like TikTok), but its full per-platform indexes are pricier ($199/mo each, or $699/mo for all platforms), with prompt-only add-ons from $50/mo. ### Do you need both Semrush and Ahrefs? Most people do not. The overlap is large, and one tool plus Google Search Console covers the core job. Teams that do serious link building sometimes keep Ahrefs for its backlink index alongside Semrush for everything else, but paying for two full subscriptions is a cost most solos should skip. ### Is there a free alternative to Semrush and Ahrefs? Google Search Console and Google Keyword Planner are free and cover your own rankings and basic keyword volumes. Both tools also offer limited free tiers, and free graders exist for specific jobs. Our [best free SEO tools](https://geotoolbox.ai/blog/best-free-seo-tools) guide covers the ones worth starting with before you pay, and our [Semrush alternatives](https://geotoolbox.ai/blog/semrush-alternatives) and [Ahrefs alternatives](https://geotoolbox.ai/blog/ahrefs-alternatives) guides compare the cheaper paid options by job. ### Which is better for a new or small website? Semrush, in most cases. In our own test it cataloged many more of the keywords our young site was picking up, so you see progress Ahrefs may not register yet. On our own new site, Semrush's keyword count grew by hundreds between our two pulls while Ahrefs's grew by seven, and its Authority Score moved from 7 to 11 (the two authority scores are separate proprietary scales, so read each against itself). If budget is the deciding factor, Ahrefs has a limited $29 Starter tier while both full plans cost about the same at entry, so lean on the free options (Search Console, Keyword Planner, and each tool's free access) first. Either way, treat the traffic estimates as directional until Search Console has enough data. ## Sources - How Accurate Are the Search Traffic Estimations in Ahrefs? (New Research) - Ahrefs (Tim Soulo), May 2022 - `ahrefs.com/blog/traffic-estimations-accuracy` - Semrush One vs SEO Toolkit plans comparison (pricing and AI engine coverage) - Semrush, August 2026 - `semrush.com/kb/1624-semrush-one-vs-seo-toolkit` - SEO & AI Search plans and pricing (replaced the SEO Classic plans page) - Semrush, checked August 3, 2026 - `semrush.com/pricing/seo-ai-search/` - AI Visibility Toolkit (features and pricing) - Semrush, August 2026 - `semrush.com/kb/1493-ai-visibility-toolkit` - Brand Radar (AI-visibility features, coverage, and pricing) - Ahrefs, August 2026 - `ahrefs.com/brand-radar` and `help.ahrefs.com/en/articles/11064852-what-is-brand-radar-and-how-to-use-it` - Adobe Completes Semrush Acquisition - Adobe press release, April 28, 2026 - `news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition` - Bot traffic directory, ranked by traffic volume - Cloudflare Radar, August 2026 - `radar.cloudflare.com/bots/directory` - AI Performance report (first-party AI-citation data for geotoolbox.ai; Bing labels the figures a sample of overall activity, sourced from Microsoft Copilots and partners) - Bing Webmaster Tools, August 2026 - `bing.com/webmasters` - Ahrefs Pricing - Ahrefs, August 2026 - `ahrefs.com/pricing` - Ahrefs vs Semrush (2026) - Which Should You Use? (10-domain referring-domain test) - Style Factory, updated July 2026 - `stylefactoryproductions.com/blog/ahrefs-vs-semrush` - Google Search Console (ground-truth click data) - Google - `search.google.com/search-console/about` - First-party account data for geotoolbox.ai and two anonymized client sites, pulled from our own Ahrefs, Semrush, Bing Webmaster Tools, and Google Search Console accounts, July 20 and August 1, 2026. --- ## How to Track Brand Mentions in AI Search > How to track brand mentions in AI search: a free method that works, the metrics that matter, how many times to run a prompt, and when a tracker is worth it. - Canonical: https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search - Published: 2026-07-20 · Updated: 2026-07-22 Brand monitoring used to mean watching the social web: alerts for your name on X, Reddit, and the news. The question now is different. When someone asks ChatGPT, Perplexity, Gemini, or Google's AI Overviews for the best tool in your category, does the answer say your name? Tracking brand mentions in AI search is how you find out, and it is more measurable than most people assume. This guide covers a free method that actually works, the handful of metrics worth watching, how many times you need to run a prompt before you can trust the result, and the point where a paid tracker starts to pay for itself. ## What Tracking Brand Mentions in AI Search Means A brand mention in AI search is any time an engine names your brand in a generated answer. It comes in three strengths, and the difference matters for what you track: - **Mentioned:** the answer names you in the text. - **Cited:** the answer names you and links to your site as a source. - **Recommended:** the answer puts you forward as the answer, not just a name in a list. The surfaces worth watching are the ones your buyers actually use: ChatGPT, Perplexity, Gemini, Claude, Google's AI Overviews and AI Mode, and Copilot among others. Each builds its answer a little differently, so a [brand mention](https://geotoolbox.ai/glossary/brand-mention) in one is no guarantee of a mention in another, and your [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) has to be read per engine. Is it even possible to track this reliably? Yes. You are not reading the model's mind, you are sampling its output: you ask the questions your buyers ask, repeatedly, and record whether you show up. The catch is that "repeatedly" is doing real work in that sentence, and it is where most tracking goes wrong. ## Why a Single Check Lies The most common mistake is to ask ChatGPT one question, screenshot the answer, and treat it as the truth. AI answers are not fixed. Ask the same question twice and you can get two different lists of brands, with nothing changed on your side. Several things move the answer between runs: - **Sampling.** The model picks each word from a probability distribution, so the same prompt takes different paths. - **Personalization.** Memory, past chats, and account context nudge what you see versus what someone else sees. - **Retrieval.** Engines that search the live web pull different pages depending on the moment. [ChatGPT may rewrite your question](https://help.openai.com/en/articles/9237897-chatgpt-search) into one or more targeted searches before it answers, and those searches are not identical every time. - **Model routing and prompt wording.** A slightly reworded prompt, or a different model version behind the same interface, changes the output. There is no fixing this; it is how the surface works. A recent study on measuring AI search, [Don't Measure Once](https://arxiv.org/abs/2604.07585), puts it plainly: answers vary across runs, prompts, and time, so one-off observations are unreliable, and visibility is best treated as a distribution rather than a single result. The practical consequence is simple. A mention is not a yes or no. It is a **rate**, and a rate needs more than one measurement. ## The Metrics That Matter Once you accept that you are measuring a rate, the metric set gets clear. You do not need all of these every week, but you should know what each one answers.
MetricWhat it answersHow to read it
Mention rateOut of N runs of a prompt, how often are you named?Your core number. Track it per prompt and per engine, not as one blended figure.
Share of voiceYour share of all brand appearances: your mentions divided by every tracked brand's mentions.Context for the mention rate. Your 30% means more when the strongest rival also sits near 30% than when one competitor is named in almost every answer.
Citation rateHow often are you named with a link to your site?Stronger than a mention: a citation can send referral traffic and shows the answer drew on your page as a source.
Position in answerAre you the recommendation, or the last name in a list?First-named brands get disproportionate attention. Being present is not the same as being prominent.
Sentiment and accuracyWhen you are named, is the framing positive, and separately, is it factually right?Score them as two columns. A confident wrong description is worse than absence, and it is fixable.
The distinction people trip over most is [mention versus citation](https://geotoolbox.ai/glossary/ai-citation). A mention is your name in the text; a citation is your name plus a link. Both matter, but they move for different reasons, so track them separately rather than folding them into one score. For a deeper look at the competitive side of this, our guide to [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) walks through how to measure your slice of the answers. ## How Many Times to Run a Prompt This is the question nobody answers, and it is the one that decides whether your tracking means anything. Most guides tell you to "sample more" and leave it there. Here is the actual math, because it changes how you read every number you collect. A mention rate is a proportion measured from a sample, exactly like a poll. And like a poll, a small sample has a wide margin of error. If you run a prompt 10 times and get named 3 times, your measured rate is 30%, but the true rate could plausibly sit anywhere in a very wide band. That is why a brand can show a 40% rate one week and 25% the next without changing a single page. Nothing changed on the site. The sample was simply too small to separate a real move from noise. The table below shows roughly how tight your number is at different sample sizes, for a rate near 30% (the margin is widest near 50% and narrower toward the extremes):
Runs per promptRough 95% marginWhat that's good for
10about ±28 pointsA gut check only. "Do I ever show up?" Not a number to report.
30about ±16 pointsSpotting large gaps against competitors.
100about ±9 pointsPinning one period's rate to about ±9 points.
~385about ±5 pointsPrecise benchmarking (this size also covers the worst case near 50%). Rarely worth doing by hand.
Two rules fall out of this. First, **fewer prompts means more runs each** - if you only track five prompts, each one carries a lot of weight, so run them more. Second, **do not react to a move smaller than your margin.** If each week's rate carries a margin of 16 points, a jump from 30% to 38% is not distinguishable from sampling noise. Comparing two periods is actually harder than measuring one, since both readings carry error, so treat small week-over-week moves as inconclusive rather than real. Judging visibility from a handful of prompts checked once or twice is like forecasting the week from one glance out the window. Two honest caveats. These margins are rough approximations, and at very small samples like 10 runs they are rougher still, so treat the low end as "directional" rather than exact. And runs fired minutes apart share the same live retrieval, so they are not fully independent; spreading runs across a few days gives a truer picture than hammering a prompt in one sitting. You do not need a statistics degree to apply this. Pick a run count that matches how precise you need to be, keep it consistent, and treat every rate as a range rather than a point. ## The Free Way to Track Brand Mentions You can run a real tracking program with a spreadsheet and an hour a week before you pay for anything. Here is the workflow. **1. Build a prompt set from unbranded questions.** The single biggest error is tracking branded prompts like "does Acme have good reviews?" The engine will almost always confirm you exist, which tells you nothing. Track the unbranded questions buyers actually type before they know you: category queries ("best AI visibility tool") and recommendation queries ("what should a small SEO team use to track AI mentions?"). Keep competitor-comparison prompts ("[rival] vs the alternatives") in a separate bucket, since naming a rival measures whether you enter the conversation, which is a different question from unaided category visibility. Aim for 10 to 20 prompts, and put your effort behind the handful that matter most so your headline rate reflects the questions buyers ask most, not an even average across trivia. **2. Run each prompt across the engines, several times.** Use fresh or incognito sessions so past chats do not color the answer, and run each prompt the number of times your target precision demands from the section above. Two engines fight this: Google's AI Overviews vary by location and do not fire on every query, and a logged-out ChatGPT can route to a lighter model. Fix your location, note when AI Overviews simply does not appear, and keep your account state consistent so the only thing changing is the model. **3. Log every run in one place.** A simple sheet with a row per run: prompt, engine, run number, mentioned (yes/no), cited (yes/no), accurate (yes/no), competitors named, and a note on [brand sentiment](https://geotoolbox.ai/blog/ai-brand-sentiment). This is the artifact almost every guide skips, and it is the whole game. Without a log you have impressions; with one you have data. **4. Compute your rates.** For each prompt and engine, divide mentions by runs to get your mention rate, and count how often each competitor appears for your share of voice. Now you have a baseline you can move. Be honest about what the free version buys you. By hand you are not going to hit 100 runs across 20 prompts, so do not pretend the spreadsheet gives you week-over-week precision. At 10 to 15 runs on your top 5 to 8 prompts, checked monthly, it answers the questions that matter early: do I show up at all, and who beats me by a wide margin? That is a real, useful read, and it is the honest ceiling of manual tracking. The paid tools exist to buy back the precision and the hours, not to make the free method fake.
![The manual AI mention-tracking loop: build a prompt set, run it across engines multiple times, log results, compute mention rate and share of voice, then act on the gaps.](/blog/track-brand-mentions-in-ai-search/manual-ai-mention-tracking-loop.png)
The manual loop: an unbranded prompt set, run repeatedly, logged, turned into rates, then acted on - and repeated.
To speed up the checking itself, a free scanner does the multi-engine legwork for you. Our free [AI Visibility Checker](https://geotoolbox.ai/features/geo-scan) runs one prompt across the eight major engines in a single pass and marks whether each one cites you, mentions you, or recommends you, with a free export you can paste straight into your log. It is a single scan rather than continuous monitoring, which is exactly what you want when you are still building the habit for free. ## Track Every Engine, Not Just ChatGPT Checking only ChatGPT is the second big blind spot. The engines disagree, often sharply. You can be the top recommendation in one and absent from another for the exact same question, because each pulls from different sources and, for the ones that search live, [assembles the answer from different retrieved pages](https://developers.google.com/search/docs/appearance/ai-features). Google's AI Overviews, for example, may break a single question into several sub-queries through a fan-out step, then draw citations from the organic index. ChatGPT may use its own search partners and rewrite your query before answering. Perplexity, Gemini, and Claude each have their own retrieval and their own preferences. The result is that a mention rate is only meaningful per engine, and one engine's picture can quietly mislead you about the others. Our breakdown of [how ChatGPT cites sources](https://geotoolbox.ai/blog/chatgpt-citations) shows why the same brand can win one engine and lose the next. You do not have to track all of them equally. Weight your effort toward the engines your buyers actually use, and check the rest less often. But check them. A program that watches one engine is measuring one slice and calling it the whole. ## When a Tracking Tool Is Worth It The free method has a ceiling. Twenty prompts, across five engines, run thirty times each, week after week during an active push, is three thousand checks logged by hand. That is the point where a tool stops being a luxury and starts saving real time. The trap to avoid is the dashboard that shows you the problem and stops there. Plenty of tools will tell you your mention rate is low and leave you to figure out what to do about it, which is why marketers have started to [question whether expensive AI visibility tools earn their price](https://digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/). What a good tool actually needs: continuous tracking across the engines, share of voice against named competitors, a clear week-over-week delta so you can see moves against the noise, and a path from the number to the fix. It should also help you close the loop on traffic. AI answers that link to you send referral visits, and you can see them in GA4 by watching for referrers like `chatgpt.com`, `perplexity.ai`, and `gemini.google.com`. Treat that traffic as the visible tip only: most AI answers resolve without a click, so referral sessions undercount how often you actually influence a buyer. That undercount is the reason to track the mention rate directly rather than inferring your AI presence from clicks alone. Across the sites we scan for AI visibility, the most common problem is not the score itself. It is that teams are flying blind: a single engine, a handful of branded prompts, checked once, mistaken for the full read. That is the gap our [AI Brand Monitoring dashboard](https://geotoolbox.ai/features/domain-overview) is built to close, tracking citation share and share of voice across all eight engines over time with a week-over-week delta. Like any tracker, it samples rather than delivering a fixed verdict, so the same run-count discipline applies. If you would rather compare the field first, our roundup of the [best AI visibility tools](https://geotoolbox.ai/blog/best-ai-visibility-tools) lays out the options. ## What to Do When You Find a Gap Tracking is only half the work. The question that follows every low mention rate is the same: the engine recommends my competitor and not me, so why, and what now? The why is usually not a secret algorithm. Generative engines build answers from the third-party sources they trust and retrieve, and the research behind [generative engine optimization](https://arxiv.org/abs/2311.09735) found that how content is written and cited materially changes whether an engine surfaces it. Your competitor is named clearly on the pages the engine pulls from, and you are missing, thin, or described with stale facts. For the engines that retrieve live sources, the answer largely repeats what those sources say rather than judging your product directly; base-model answers move more slowly but still lean on how widely and consistently you appear in training data. So the fix follows the log. Take the prompts you lose and look at which sources the engine cites in those answers. Then go earn a presence there: the review sites, the listicles, the community threads, and the industry media that keep appearing. One mention rarely moves anything; what shifts an answer is consensus, the same brand described consistently across several independent sources the engine already trusts. Make sure your own facts line up across the web so the engine has a clear, repeated signal, which is the core of strong [entity SEO](https://geotoolbox.ai/glossary/entity-seo). Then re-track the same prompts and watch the rate settle over the next few weeks. The full earning-side playbook, from crawler access to citable structure, is our guide to [how to get cited by AI](https://geotoolbox.ai/blog/how-to-get-cited-by-ai). A worked example: say your mention rate for "best AI mention tracker" is 10% in Perplexity, and the answers keep citing the same three comparison articles. You do not have to move the model. You have to get accurately represented in those three articles, then confirm the lift over your next measurement window. That is the whole loop, and it is why tracking and fixing belong in the same workflow rather than two separate tools. For a structured version of this, our [AI visibility audit](https://geotoolbox.ai/blog/ai-visibility-audit) walks through it step by step. ## Frequently Asked Questions ### Is it possible to track brand mentions in AI search? Yes. You cannot see inside the model, but you can sample its output: run the prompts your buyers ask, repeatedly, across each engine, and record whether your brand appears. Because answers vary run to run, you track a mention rate over many runs rather than reading a single response. That is measurable with a spreadsheet or a tool. ### How often should I check whether AI engines mention my brand? Match the cadence to how fast your category moves and how much you are actively working on it. Monthly is enough for a stable baseline; weekly makes sense when you are running a campaign to earn mentions. More important than frequency is consistency: same prompts, same run count, so the numbers are comparable over time. ### What's the difference between a brand mention and a citation in AI answers? A mention is your name in the generated text. A citation is your name plus a link to your site as a source. A citation is the stronger signal because it can send referral traffic and shows the answer drew on your page as a source. Track them as two separate metrics, since they move for different reasons. ### Are there free tools to track brand mentions in AI search? Yes, for the checking step. Free scanners, including our own [AI Visibility Checker](https://geotoolbox.ai/features/geo-scan), run a prompt across the major engines and show whether you are cited, mentioned, or recommended. They are single scans rather than continuous monitors, so pair them with a logging spreadsheet to build a rate over time. Ongoing, automated tracking across many prompts is where paid tools come in. ### Why do the AI answers change every time I check? Because the models sample their output and, for the ones that search live, retrieve different pages each time. Personalization, prompt wording, and model routing add more variation. This is normal and unavoidable, which is exactly why a single check is unreliable and you measure a rate across repeated runs instead. ### Is tracking ChatGPT enough, or do I need the other engines? Not enough on its own. Engines pull from different sources and can disagree sharply, so you can lead in ChatGPT and be absent in Gemini for the same question. Track the engines your buyers actually use, and check the others less often rather than ignoring them. ## Start Tracking This Week The reason to start now is not that AI search is coming, it is that it already answers without a click. Bain reports that [about 60% of searches now end without the user moving on to another site](https://www.bain.com/insights/goodbye-clicks-hello-ai-zero-click-search-redefines-marketing/), a shift it ties to the rise of AI summaries. When the answer is the destination, whether it names you is what you need to know. You do not need budget to find out where you stand. Build an unbranded prompt set, run it across the engines a few times, and log what comes back. Run a first pass through the free [AI Visibility Checker](https://geotoolbox.ai/features/geo-scan) to see your eight-engine picture in minutes, and when the spreadsheet gets heavier than the insight, move the whole thing into the [AI Brand Monitoring dashboard](https://geotoolbox.ai/features/domain-overview) and let it keep the rate for you. ## Sources - Goodbye Clicks, Hello AI: Zero-Click Search Redefines Marketing - Bain & Company, 2025 - `bain.com/insights/goodbye-clicks-hello-ai-zero-click-search-redefines-marketing/` - Don't Measure Once: Measuring Visibility in AI Search (GEO) - arXiv, Schulte, Bleeker & Kaufmann, April 2026 - `arxiv.org/abs/2604.07585` - ChatGPT Search - OpenAI Help Center - `help.openai.com/en/articles/9237897-chatgpt-search` - AI Features and Your Website - Google Search Central - `developers.google.com/search/docs/appearance/ai-features` - GEO: Generative Engine Optimization - arXiv, Aggarwal et al. (ACM SIGKDD 2024) - `arxiv.org/abs/2311.09735` - Marketers question expensive AI visibility tools as inconsistent results fuel skepticism - Digiday, Kimeko McCoy, May 2026 - `digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/` --- ## What Is an XML Sitemap? Format, Limits, and How to Find One > What an XML sitemap really is, which tags Google still uses in 2026, how to find and validate one, and whether it helps rankings or AI search. - Canonical: https://geotoolbox.ai/blog/xml-sitemap - Published: 2026-07-20 · Updated: 2026-07-20 An XML sitemap is a file that lists the URLs on your site so search engines can find them. That is the whole job. Most of what gets written about sitemaps oversells it from there, so this guide sticks to what is actually true in 2026: what the file is, what each tag does (and which ones Google ignores), how to find and validate one, and the real answer to whether a sitemap helps your rankings or your visibility in AI search. Paste any domain into the free [Sitemap URL Extractor & Validator](/tools/sitemap-extractor) and it finds the sitemap, pulls every URL out, and flags what is broken. No sign-up. ## What an XML Sitemap Actually Does An XML sitemap is a machine-readable list of the URLs you want search engines to know about, saved as an XML file and usually served at `/sitemap.xml`. It exists so crawlers can discover your pages without relying entirely on following links. That is the useful part. Here is the part most guides skip: a sitemap is a **discovery aid, not a command**. Google's own documentation says a sitemap "doesn't guarantee that all the items in your sitemap will be crawled and indexed." Listing a URL is a hint that the page exists and is worth looking at. It does not force Google to crawl it, index it, or rank it. So a sitemap earns its keep in specific situations. New sites with few inbound links get discovered faster, because there are not many other paths to the pages yet. Large sites get more complete coverage, because a crawler following links alone can miss deep or poorly-linked pages. The same is true for JavaScript-heavy or single-page sites, where a crawler may not follow client-rendered navigation, so the sitemap becomes the main way pages get found. For a small, well-linked site, the file often changes nothing, which is a point we come back to. What a sitemap does **not** do is worth stating plainly, because the myths are persistent. It is not a ranking factor. It does not fix "not indexed" pages on its own. And it does not push a page into the index that Google has decided not to keep. When a page in your sitemap stays unindexed, the URL was still found; the problem is downstream, usually page quality, duplication, or internal linking. ## What an XML Sitemap Looks Like At its core, a sitemap is a `` element wrapping one `` block per page. Each block needs a `` (the address) and can carry a few optional hints. Here is a minimal, valid example: ```xml https://example.com/ 2026-07-18 https://example.com/blog/xml-sitemap 2026-07-20 ``` A few rules from the [sitemaps.org protocol](https://www.sitemaps.org/protocol.html) that trip people up. The file must be UTF-8 encoded. URLs must be fully-qualified and absolute (`https://example.com/page`, not `/page`). And any special characters in a URL have to be entity-escaped: an ampersand becomes `&`, a `<` becomes `<`, and so on. An unescaped `&` is a common reason a hand-built sitemap fails to parse. In practice you rarely write this file by hand: a CMS or an SEO plugin generates it for you, and standalone generators exist for static sites. Knowing the shape still matters, because it is how you spot when a generated one is wrong. The `xmlns` attribute on `` is the sitemap namespace and is required. It is not decoration; it tells parsers which schema the file follows. Every entry after that is just a `` with its `` and, optionally, the four hint tags below. The standard `` format above covers almost every site. There are also specialized sitemaps for particular content: image, video, and Google News sitemaps (each adds its own XML namespace), plus hreflang sitemaps that declare alternate-language versions of a page for international sites. Most sites never need them, and they are beyond this guide's scope, but it helps to know they exist if you run a news, media-heavy, or multilingual site.
TagRequired?What it is forDoes Google use it?
<loc>YesThe absolute URL of the pageYes, this is the point of the file
<lastmod>NoWhen the page last changed, in W3C datetime formatYes, but only if it is consistently accurate
<changefreq>NoA hint at how often the page changesNo, ignored
<priority>NoRelative importance, 0.0 to 1.0No, ignored
## The Four Tags, and Which Ones Still Matter The `changefreq` and `priority` tags are not worth tuning, and Google says so directly. Its [sitemap documentation](https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap) states that "Google ignores `` and `` values." Setting every page to `priority` 1.0 does nothing. Declaring `changefreq` as `daily` does not make Googlebot visit daily. Other engines may read them, but none give them enough weight to matter. If your generator emits them, it is harmless; just do not spend time on them. That leaves ``, the one metadata tag that actually earns attention, precisely because Google does use it, with a catch. Google treats `` as a signal only when it is **demonstrably accurate**, meaning the date reflects a real, significant change to the page content. When the dates are trustworthy, they help Google prioritize what to re-crawl. The catch is what happens when they are not. Plenty of CMS setups stamp today's date on every URL in the sitemap on every rebuild, so the whole file claims everything changed this morning. Google notices that pattern, concludes the `` values are noise, and stops trusting them for your site. An inflated `` is no better than none, because Google learns to ignore it and you lose the signal you were trying to send. If your platform bumps every date automatically, that is a real thing to check, and it shows up clearly when you pull the full URL list and see identical dates down the whole column. So the 2026 rule for tags is short. Get `` right, keep `` honest or leave it off, and ignore `` and `` completely. ## Big Sites: Index Files, Gzip, and the Limits A single sitemap file is capped. The protocol allows a maximum of **50,000 URLs** or **50MB uncompressed** per file, whichever you hit first (that byte ceiling is exactly 52,428,800 bytes). A 200,000-page site therefore cannot use one flat file. It splits into several, grouped under a **sitemap index**. A sitemap index is a file whose entries point at other sitemap files instead of at pages. It uses a `` root, and each `` entry has a `` for the child file and an optional `` for when that child last changed: ```xml https://example.com/sitemap-posts.xml 2026-07-20 https://example.com/sitemap-products.xml.gz 2026-07-19 ``` This is the normal setup on any site past a few hundred pages. WordPress, Shopify, Yoast and Rank Math all publish an index by default. It matters because it is where a lot of tools quietly fail: paste an index into a basic extractor and you get back a list of five child sitemap URLs instead of your five thousand pages. The file did what it was told; it just did not do the job you wanted.
LimitValueApplies to
URLs per file50,000A single sitemap
Uncompressed file size50MB (52,428,800 bytes)A single sitemap or index
Sitemaps per index50,000A sitemap index file
<loc> URL lengthUnder 2,048 charactersEach URL
Two more details save headaches. Sitemaps can be **gzip-compressed** and served as `.xml.gz`, which helps on large files, but the 50MB limit is measured on the **uncompressed** size, not the `.gz`. And all the URLs in a sitemap must sit on the **same host** as the sitemap itself; crawlers ignore off-host URLs unless you have authorized cross-submission, either by verifying both sites in Search Console or by referencing the sitemap from the other host's `robots.txt`. ## How to Find the Sitemap of Any Website Most people start by typing `/sitemap.xml` and give up if it 404s. That is the wrong first move, because a site can put its sitemap anywhere and tell you exactly where. The order that actually works: 1. **Read `robots.txt` first.** Fetch `example.com/robots.txt` and look for `Sitemap:` lines. A site can declare one or several, pointing anywhere, including a different host. This is the site telling you directly, so it beats guessing. 2. **Probe the common paths.** If robots.txt is silent, try `/sitemap.xml`, `/sitemap_index.xml`, `/sitemap-index.xml`, and `/wp-sitemap.xml`. That last one is WordPress core since version 5.5 and is an easily missed location. 3. **Check Search Console.** If it is your own site, the Sitemaps report shows what you actually submitted, which is often a different set from what robots.txt advertises. One thing not to bother with: a `site:` search does not reliably surface a sitemap. Sitemaps are XML files, not pages, so an empty `site:` result tells you nothing about whether one exists. That bare-domain-to-URL-list step is exactly what our free [sitemap extractor](/tools/sitemap-extractor) automates: give it just a domain, and it reads robots.txt, probes the common paths including `wp-sitemap.xml`, walks nested index files (the ones a basic extractor stops at), decompresses `.xml.gz`, and hands you every URL, then tells you which method found the file. ## Validating a Sitemap: Two Things "Valid" Means "Is my sitemap valid?" is really two questions, and most tools only answer the first one. The first is **spec conformance**: does the XML match the schema? Required ``, well-formed dates, `priority` inside 0.0 to 1.0, a recognized `changefreq`, escaped ampersands, under 50,000 URLs and 50MB. This is the easy part, and a syntax validator or Search Console will catch most of it. The second is **truthfulness**, and it is where the real damage hides: are those URLs telling the truth? A sitemap can be flawless XML and still be useless if every URL in it 404s, redirects, points at a `noindex` page, or sits on a host you did not intend. Search engines treat a sitemap full of non-200 URLs as a low-quality signal about the whole file. So the questions that matter are: do the URLs still resolve, are they on the right host, and does anything else on the site contradict them? That contradiction check is the one people skip, because it means cross-referencing the sitemap against `robots.txt` and against live HTTP status codes, not just reading the XML. The same free tool runs both passes in one go: it scores the file against the sitemaps.org spec, cross-checks `robots.txt`, and lets you fire live status checks at the URLs, so you catch the "valid but lying" cases before you submit. The most useful time to run it is right before submitting to Search Console, so errors never get logged against your property in the first place. ## Sitemaps and robots.txt: The Conflict That Blocks Pages The sitemap and [robots.txt](/glossary/robots-txt) are really one job seen from two sides. The sitemap says what exists; robots.txt says what may be fetched. The damage happens where they disagree. The classic own-goal is a URL that appears in your sitemap while being blocked by a `Disallow` rule in robots.txt. You are asking Google to crawl a page and, in the same breath, telling it not to fetch that page. In Search Console this shows up as "Blocked by robots.txt" or, if Google indexed the URL anyway on the strength of links pointing at it, the "Indexed, though blocked by robots.txt" warning. That second warning confuses people because it looks contradictory, so be precise about what robots.txt actually controls. Per Google, "a page that's disallowed in robots.txt can still be indexed if linked to from other sites." Robots.txt governs **crawling**, not **indexing**. It is, in Google's words, "not a mechanism for keeping a web page out of Google." That distinction dictates the fix, and the fix is the opposite of most people's instinct. If the page **should** be indexed, remove the `Disallow` rule that blocks it, and make sure the URL belongs in the sitemap. If the page should **not** be indexed, do not leave it blocked; unblock it in robots.txt and add a `noindex` tag instead, because Google has to be able to crawl the page to see the `noindex`. Then take that URL out of the sitemap. A URL that is both submitted and disallowed is the contradiction to eliminate either way. If you are not sure which rule is doing the blocking, the free [robots.txt tester](/tools/robots-txt-tester) names the exact line. ## How to Submit Your Sitemap to Google There are two methods that still work, and you should use both. First, declare the sitemap in `robots.txt` with a `Sitemap:` line pointing at the absolute URL (you can list more than one). This is passive discovery for any crawler that reads the file. Second, submit it in **Google Search Console** under Indexing, then Sitemaps. This is the one that gives you feedback: fetch status, how many URLs Google read, and the errors it found. If you have read an old tutorial, you may have seen a third method, "pinging" a special Google URL to announce an update. That is gone. Google [deprecated the sitemaps ping endpoint](https://developers.google.com/search/blog/2023/06/sitemaps-lastmod-ping) on June 26, 2023, and switched it off roughly six months later; requests to it now return a 404. Google's stated reason was that the unauthenticated pings were rarely useful and mostly generated spam. So do not wire a ping into your publishing pipeline, and if you have one, remove it. The replacement for "tell Google I changed something" is the ordinary path: keep an accurate `` and let Search Console and normal recrawling do the rest. Note that IndexNow, which Bing and some other engines support, is a separate mechanism that still exists; Google's ping going away does not affect it. ## Do Sitemaps Help Rankings or AI Search? No, and a lot of guides fudge this. A sitemap is not a ranking factor. You will see claims that submitting a sitemap "boosts rankings" or "signals trust and authority." Google's position is the opposite: a sitemap is a discovery hint, and the protocol itself notes that `priority` values are "not likely to influence the position of your URLs." A sitemap can help a page get **found**, which is a prerequisite for ranking, but finding is not ranking. If a page is already discoverable through internal links, the sitemap adds nothing to how it ranks. The AI-search version of the question deserves the same candor. A sitemap is not a citation signal, and no AI provider documents it as one. The one carve-out is news publishing: a Google News sitemap speeds timely inclusion in Google News (it is not required, but it helps), and that content can surface in AI answers, so for a news site it does have an indirect effect. For everyone else, it does not. What decides whether ChatGPT, Perplexity, or Google's AI answers can use your pages is **reachability**: whether their crawlers are allowed to fetch you and whether the page renders into readable content. That is a robots.txt and rendering question, not a sitemap one. A perfect sitemap in front of pages that block [AI crawlers](/blog/ai-crawlers) does nothing for AI visibility.
![Comparison of sitemap.xml, robots.txt, and llms.txt: what each file declares, whether it gates crawler access, and whether it is a ranking or citation signal.](/blog/xml-sitemap/sitemap-robots-llms-txt-compared.png)
Three files, three jobs. Only robots.txt actually gates whether a crawler reaches your pages.
This is where the files stack up, and keeping them straight helps. The sitemap says what exists. robots.txt decides who is allowed to fetch it. [llms.txt](/glossary/llms-txt), the newer proposal, tries to hand models a curated map, though the measured evidence that AI engines read it is thin. Of these, the one that actually gates AI access today is robots.txt. When we audit sites for AI visibility, a recurring blocker we see is not a missing or malformed sitemap; it is a robots.txt rule quietly disallowing an AI crawler the owner never meant to block. If you want to check that side, our free [AI Crawler Checker](/tools/ai-crawler-checker) shows which of 34 AI crawlers your robots.txt allows or blocks, with the exact line to change. ## Frequently Asked Questions ### Do I need a sitemap for a small website? Often not. Google says a site that is "small" (about 500 pages or fewer) and thoroughly linked internally, so Googlebot can reach every important page from the home page, may not need a sitemap at all. If that describes you, a sitemap will not hurt, but it is not the thing holding your indexing back. It matters most for large sites, brand-new sites with few inbound links, and deep pages that internal linking misses. ### How often should an XML sitemap update? Automatically, whenever your content changes, which is what any CMS or sitemap plugin handles for you. There is no schedule to set and no benefit to regenerating it on a timer. The one thing to keep honest is ``: it should reflect real changes, not today's date stamped on every URL, or Google learns to ignore it. ### What is the difference between a sitemap and a sitemap index? A sitemap lists page URLs. A sitemap index lists other sitemap files. Once a site passes 50,000 URLs (or 50MB), it must split its pages across multiple sitemaps and publish an index that points at all of them. In practice CMSs do this long before the limit, so most sites above a few hundred pages ship an index by default, which is why an extractor has to follow the children to give you the actual pages. ### Can an XML sitemap be gzipped? Yes. Serving it as `.xml.gz` is common and sensible on large files. One caveat: the 50MB size limit applies to the uncompressed file, not the compressed `.gz`, so compression saves bandwidth but does not raise the ceiling. Note that a `.xml.gz` is a gzipped file, not a gzip-encoded HTTP response, which is why some browsers and tools show you binary junk instead of reading it. ### Does an XML sitemap help my site get cited by ChatGPT or Perplexity? Only indirectly, and less than people hope. A sitemap can help your pages get discovered and indexed, which is a precondition for being used in AI answers, but it is not a citation signal and no AI engine documents it as one. What decides whether an AI crawler can use your content is whether it is allowed to fetch the page and whether the page renders cleanly, which is a robots.txt and rendering matter, not a sitemap one. ### What is the difference between an XML sitemap and an HTML sitemap? An XML sitemap is for machines: a structured URL list search engines read to discover pages. An HTML sitemap is a page for humans, a browsable directory of links. They solve different problems. Good internal linking usually does the HTML sitemap's job better; the XML sitemap is the one that matters for crawl discovery. ## The Bottom Line An XML sitemap is plumbing, not strategy. Get it complete and accurate, keep `` honest, make sure nothing in it is contradicted by `robots.txt`, and then stop tuning tags that Google ignores. Most of the effort people spend "optimizing" a sitemap is better spent making sure the URLs in it resolve and are allowed to be crawled. If you want to see the state of any sitemap in one pass, that is what we built the free [Sitemap URL Extractor & Validator](/tools/sitemap-extractor) for: paste a domain, get every URL out, and see what is broken, off-host, or blocked by your own robots.txt, before Google does. It is free, with no sign-up. ## Sources - Build and Submit a Sitemap - Google Search Central (limits, ignored tags, lastmod, submission, cross-submission) - `developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap` - Sitemaps overview / Do you need a sitemap? - Google Search Central (no crawl or index guarantee; the ~500-page threshold) - `developers.google.com/search/docs/crawling-indexing/sitemaps/overview` - Sitemaps XML format - sitemaps.org Protocol (tags, limits, entity escaping, sitemap index, gzip, same-host rule) - `sitemaps.org/protocol.html` - Sitemaps ping endpoint is going away - Google Search Central Blog (deprecation announced June 26, 2023) - `developers.google.com/search/blog/2023/06/sitemaps-lastmod-ping` - Introduction to robots.txt - Google Search Central (robots.txt controls crawling, not indexing; use noindex to keep pages out) - `developers.google.com/search/docs/crawling-indexing/robots/intro` --- ## Kimi API Pricing August 2026: Cost, Keys, and Limits > What the Kimi API costs in 2026: per-model pricing, the $1 minimum, rate-limit tiers, and the endpoint split that quietly returns a 401 on a valid key. - Canonical: https://geotoolbox.ai/blog/kimi-api-pricing - Published: 2026-07-18 · Updated: 2026-08-08 The Kimi API costs $3.00 per million input tokens and $15.00 per million output tokens on kimi-k3, drops to $0.95 and $4.00 on the coding model, and requires a $1 minimum top-up before it will answer a single request. There is no free tier. Those numbers are easy to find and easy to get wrong. We asked five AI engines what the Kimi API costs. Three produced confident tables whose highest input quote sat 73% above the lowest, because they were reading resellers instead of Moonshot's own docs. Everything below comes from the pricing pages themselves, checked July 18, 2026. Three things catch developers out: the endpoint split that returns a silent 401, the 429 that really means your balance is empty, and why prompt caching will not make K3 as cheap as you are hoping. ## How Much Does the Kimi API Cost? **Kimi K3, the current flagship, costs $3.00 per million input tokens and $15.00 per million output tokens, with cached input dropping to $0.30.** Every model on the platform is billed in US dollars, and every price below comes from Moonshot's own pricing docs, checked directly. The column most pricing roundups omit is the middle one. Moonshot charges a tenth of list for cached input on K3 and roughly a fifth to a sixth on the other K2-line models. The legacy `moonshot-v1` family gets no cache discount at all, so what you pay depends on your prompt repetition and on which model you picked.
Model IDInput (cache miss)Input (cache hit)OutputContextStatus
kimi-k3$3.00$0.30$15.001,048,576Current flagship
kimi-k2.7-code-highspeed$1.90$0.38$8.00262,144Faster variant
kimi-k2.7-code$0.95$0.19$4.00262,144Coding workhorse
kimi-k2.6$0.95$0.16$4.00262,144General purpose
kimi-k2.5$0.60$0.10$3.00262,144Sunsets Aug 31, 2026
moonshot-v1-128k$2.00None$5.00131,072Sunsets Aug 31, 2026
moonshot-v1-32k$1.00None$3.0032,768Sunsets Aug 31, 2026
moonshot-v1-8k$0.20None$2.008,192Sunsets Aug 31, 2026
Prices are per million tokens and exclude tax, which Moonshot calculates at checkout based on your jurisdiction. The K3 row and the legacy rows deserve a closer look. The first is K3. Kimi built its reputation as [the cheap open-weight option](https://geotoolbox.ai/blog/what-is-kimi-ai), and K3 breaks that. At $15.00 per million output tokens it costs 3.75 times what kimi-k2.7-code charges, which puts it in the same conversation as the frontier models it was built to compete with rather than the budget tier developers came to Kimi for. The second is the legacy row. The `moonshot-v1` family is still on sale, and it is a trap: `moonshot-v1-128k` costs more per input token than `kimi-k2.6` while giving you half the context and no cache discount at all. If you are running an integration built in 2025 and nobody has revisited the model string since, you are paying more for less. Both families retire on the same date. If you want the full picture on what K3 actually is before deciding whether the flagship rate is justified, our [Kimi K3 explainer](https://geotoolbox.ai/blog/what-is-kimi-k3) covers the specs, the benchmarks worth trusting, and whether you can realistically self-host it. ## Is There a Free Kimi API Tier? **No. You need to put money in before the API will answer you at all, and the minimum is $1.** There is no permanent free tier and no trial credit. This is the most common misconception about Kimi, and the search data shows it: most questions Google surfaces about the Kimi API are some version of "is it free?" The confusion is understandable, because the consumer Kimi app at kimi.com does have a free tier, and its paid plans are metered in credits rather than tokens. We cover that side separately in our [Kimi pricing](https://geotoolbox.ai/blog/kimi-pricing) guide. The API is a separate product with separate billing. A top-up buys access and a rung on the rate-limit ladder. Recharge amounts are non-refundable, so treat the first payment as a commitment rather than a trial. Spend $5 cumulatively and Moonshot issues a $5 voucher, which is the closest thing to free credit on offer. One catch before you plan around it: vouchers do not count toward tier progression. Only real money moves you up. Moonshot also ran a K3 launch top-up rebate over the summer: a voucher worth 10% to 30% of a top-up at the $20, $100, $300, and $1,000 thresholds, once per organization, capped at $4,000, and excluding both the Kimi app and Kimi Code. Per the promo page as captured on July 18, 2026 (since taken down), it was set to run into mid-August, so do not count on it; the standing $5 voucher above is the credit you can actually rely on. If a similar rebate returns, the mechanics are worth knowing: it was calculated from the highest-value transaction made on the day of your first top-up, so a $20 test and a $1,000 commit on the same day earned 30% on the $1,000, a commit that landed the following day did not count, and vouchers expired after 90 days. ## How to Get a Kimi API Key **Create an account on Moonshot's developer platform, top up at least $1, and generate a key from the console.** The steps take about five minutes. Picking the right platform to do them on is what costs people an afternoon. Moonshot runs more than one API surface, and a key issued on one will not authenticate against another. The failure is silent: you get a bare 401 with no hint that the key itself is valid and simply pointed at the wrong door. It is a common enough mistake to show up repeatedly in integration issue trackers.
What you haveBase URLPlatformHow it bills
Key from platform.kimi.aihttps://api.moonshot.ai/v1Developer Platform (international)Per token, from a prepaid balance
Key from platform.moonshot.cnhttps://api.moonshot.cn/v1Developer Platform (China)Per token, separate balance
Key from the Kimi Code consolehttps://api.kimi.com/coding/v1Kimi CodeSubscription; ~300-1,200 requests per rolling 5-hour window by plan
![Table mapping each Kimi key source to its base URL and billing model: platform.kimi.ai to api.moonshot.ai, platform.moonshot.cn to api.moonshot.cn, and Kimi Code console keys to api.kimi.com/coding/v1.](/blog/kimi-api-pricing/kimi-api-endpoint-routing.png)
A Kimi key only works against the platform that issued it. Mixing them returns a bare 401 with no explanation.
The international and China platforms hold separate balances. Money you top up on one is not spendable on the other, and the accounts do not sync. Pick the one that matches where you are billing from and keep every key, environment variable, and base URL on that side of the line. Kimi Code looks like a cheaper way to buy the same access. It is a different product: a subscription with its own keys and its own rolling request window, aimed at coding agents instead of metered calls. It is also what the Kimi CLI authenticates against, which is why a metered platform key will not drive it. A Kimi Code subscription issues its own keys, and [its documentation](https://www.kimi.com/code/docs/en) covers both OpenAI and Anthropic protocol formats. Metered platform keys speak OpenAI's format only, which is why dropping a standard Kimi key into an Anthropic-configured agent tool fails with nothing more useful than a 401. Once you are on the right platform the rest is ordinary. The API is [OpenAI-SDK compatible](https://platform.kimi.ai/docs/introduction), so pointing an existing client at the base URL and swapping the model string is usually the whole integration: ```python from openai import OpenAI client = OpenAI( api_key=os.environ["MOONSHOT_API_KEY"], base_url="https://api.moonshot.ai/v1", ) resp = client.chat.completions.create( model="kimi-k3", messages=[{"role": "user", "content": "Hello"}], ) ``` The environment variable is `MOONSHOT_API_KEY`, not `KIMI_API_KEY`, and the host is `api.moonshot.ai`, not `api.kimi.ai`. The docs moved to Kimi branding; the SDK surface did not. **Note:** *A paid subscription to the consumer Kimi app does not include metered API credits. Moonshot bills the API Open Platform separately from the Kimi app and Kimi Code.* ## Rate Limits and the Tier System **Your rate limits are set by how much you have topped up in total, not by what you spend per month.** Moonshot runs [six tiers keyed to cumulative recharge](https://platform.kimi.ai/docs/pricing/limits), and the first step up is the one that matters.
TierCumulative top-upConcurrencyRequests per minuteTokens per minute
Tier 0$113500,000
Tier 1$10502002,000,000
Tier 2$201005003,000,000
Tier 3$1002005,0003,000,000
Tier 4$1,0004005,0004,000,000
Tier 5$3,0001,00010,0005,000,000
Going from Tier 0 to Tier 1 costs $9 and buys 50 times the concurrency. If you are evaluating Kimi with a $1 balance and concluding it is too slow for real work, you are measuring the entry tier, not the API. Tier 0 also carries the only daily token cap in the system, at 1.5 million tokens per day; every tier above it is uncapped on that axis. The tokens-per-minute ceilings above are also a published maximum rather than a guarantee: Moonshot notes that temporary adjustments may occur when cluster demand is high. Three details catch people out, and all three cost money. First, limits are enforced **per user, not per key**, so minting a second key to work around throttling does nothing. Second, they are **shared across models**, so a batch job on kimi-k2.6 eats into the same budget your K3 calls are drawing from. Third, Moonshot meters your rate limit against `max_completion_tokens`. Not the tokens you spend, the ceiling you declared. Setting that parameter to an optimistic ceiling reserves capacity you never consume, which throttles you earlier than your real usage warrants. Set it close to what you expect. Billing is unaffected either way: you are charged for the tokens actually generated, not the ceiling you reserved. ## The 429 That Means Your Balance Is Empty A dead Kimi integration and a throttled one return the same status code. When the balance hits zero, the API has been observed answering 429 Rate Limit Reached, with nothing in the status code to distinguish it from ordinary throttling, and that has a specific consequence: a retry layer that keys off the status code alone will treat a terminal billing condition as a transient one and keep backing off until it exhausts its attempts. The [openclaw project logged exactly this](https://github.com/openclaw/openclaw/issues/43447), filing it as "Kimi/Moonshot 'Rate Limit' error masks insufficient funds, causes UI lockout." It has since been fixed on their side, in a patch that classifies a Moonshot balance 429 as a billing failure rather than throttling. That fix lives in openclaw, though. It is client-side handling of an API behavior that still exists, so any tool that has not special-cased it will keep spinning. The fix is one call on the error path. On a 429, check your balance before you back off: ```python api_key = os.environ["MOONSHOT_API_KEY"] class BillingError(RuntimeError): pass def handle_429(response: httpx.Response) -> None: """Illustrative error-path logic. A production policy still needs jitter, retry limits, and a cached balance result.""" if response.status_code != 429: return try: r = httpx.get( "https://api.moonshot.ai/v1/users/me/balance", headers={"Authorization": f"Bearer {api_key}"}, timeout=5, ) r.raise_for_status() balance = float(r.json()["data"]["available_balance"]) except (httpx.HTTPError, KeyError, TypeError, ValueError): time.sleep(1) # balance check failed; treat as retryable return if balance <= 0: raise BillingError("Kimi balance exhausted, not rate limited") time.sleep(1) ``` Cache that result rather than calling it on every 429, or a retry storm will add one request per failure to an account that is already refusing them. Moonshot exposes a balance endpoint precisely so you can make this distinction. One call on the error path separates "wait and try again" from "nothing you do will help until someone tops up the account," which is the difference between a slow job and a dead one nobody notices. Moonshot documents no auto-recharge, no spend alert, and no low-balance notification, and top-ups are manual, so that check is the only signal you get, and it only fires after a call has already failed. ## What the Bill Looks Like Most Kimi API pricing coverage quotes the cache discount and stops there. That misses the distinction that matters: caching changes your bill a lot, and barely changes which model is cheaper. Here is why. Cache discounts apply to input tokens only. Output tokens are never cached, they are billed at full rate every time, and K3's output rate is 3.75 times kimi-k2.7-code's. So as your cache hit rate improves, input costs shrink toward nothing and output becomes almost the entire invoice, which is exactly the part where K3 is most expensive. Take a workload of 1 million input and 100,000 output tokens per day, a 10:1 input-to-output ratio, and run it at no caching and at an 85% hit rate. The ratio matters: agentic coding runs output-heavier, which sharpens the conclusion below, while heavy summarization runs input-heavier and softens it.
ModelCache hit rateInputOutputDaily totalOutput share
kimi-k30%$3.000$1.500$4.50033%
kimi-k385%$0.705$1.500$2.20568%
kimi-k2.7-code0%$0.950$0.400$1.35030%
kimi-k2.7-code85%$0.304$0.400$0.70457%
![Line chart of daily Kimi API cost against prompt-cache hit rate. The kimi-k3 line falls from about $4.50 at a 0% hit rate to $2.21 at 85%, and the kimi-k2.7-code line from about $1.35 to $0.70, so the ratio between them barely moves even as both lines fall.](/blog/kimi-api-pricing/kimi-cache-cost-curve.png)
Caching cuts both bills by about half. It barely moves the ratio between the two models, because output tokens are never cached.
Caching cuts the K3 bill by 51% and the k2.7-code bill by 48%. It moves the ratio between them from 3.33x to 3.13x. One caveat on that 85%: Moonshot documents no cache API, no TTL, and no keying rules, so there is nothing published to tune against; treat your hit rate as an outcome you measure rather than a number you can engineer. If you picked K3 expecting aggressive prompt caching to bring it near the coding model's price, it does not, and no achievable hit rate will. The lever works on the shrinking half of the invoice. Batch and web search change the math. **Batch inference is 40% off.** Moonshot prices [batch jobs](https://platform.kimi.ai/docs/pricing/batch) at 60% of standard rates, which is the largest single saving available on the platform. It comes with a restriction the pricing page makes plain: the batch table covers kimi-k2.7-code, kimi-k2.6, and kimi-k2.5. K3 is not on it. Batch jobs are also exempt from the real-time concurrency limits above, which matters more than the discount if you are sitting at Tier 0 with a single concurrent request. If your workload tolerates asynchronous processing, that is a strong argument for staying on the K2 line. **Web search is billed per call.** The [`$web_search` tool](https://platform.kimi.ai/docs/pricing/tools) costs $0.005 per successful call on top of tokens, and the search results it injects are billed as input on the following call, which on K3 routinely costs more than the call fee itself. Moonshot currently carries a warning on that page saying the feature is being updated, that the documentation is outdated, and that it does not recommend using the functionality in the near term. Treat both the price and the behavior as provisional. That note was live as of July 18, 2026. None of this is unique to Kimi. The same trap catches teams running frontier models anywhere, which is why we wrote about [how token costs add up](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) in agentic workflows, and why the comparison to run is against [Gemini API pricing](https://geotoolbox.ai/blog/gemini-api-pricing) on your own traffic shape, not on a list price. ## Direct, OpenRouter, or Kimi Code? Moonshot sells the same weights through resellers, and the arbitrage runs in an unintuitive direction. Buying direct is not automatically cheapest, and on one model it is measurably the most expensive.
RouteBest forHow it pricesWatch for
Moonshot directK3, production workloads, batch jobsPer token, prepaid balanceWeChat Pay and Alipay only for individuals
OpenRouterK2.5, multi-provider fallback, card paymentPer token, pass-through plus marginUndercuts direct on K2.5, exact parity on K3
Kimi CodeCoding agents, predictable monthly spendSubscription, rolling 5-hour quotaNot metered credit; separate keys and limits
[OpenRouter lists kimi-k2.5](https://openrouter.ai/moonshotai/kimi-k2.5) at $0.375 input and $2.025 output per million tokens. Moonshot charges $0.60 and $3.00 for the same model. That is 37.5% cheaper on input and 32.5% cheaper on output, through a reseller, for the same advertised model family. On K3 the same comparison lands at exact parity, $3.00 and $15.00 either way, so the discount is a K2.5 artifact rather than a standing rule. There is a wrinkle that makes it more than trivia. K2.5 sunsets on August 31, 2026 and Moonshot's [model list](https://platform.kimi.ai/docs/models) already marks it "no longer available to new users", so a new first-party account may no longer be able to select it, while OpenRouter still listed it when we checked. That is a strange place for a vendor's own economics to end up, and it is a reason to check before assuming the first-party option wins on price. Do not build on it, though: OpenRouter is a pass-through, so the K2.5 route very likely closes at the same August 31 sunset, and the nearest replacement moves cache-miss input from $0.60 to $0.95. Payment method may decide this for you regardless. Moonshot documents WeChat Pay and Alipay for individual accounts and never mentions cards. Invoices are issued by Beijing Moonshot AI Technology Co., Ltd. with 6% Chinese VAT, which sits awkwardly beside the pricing pages' statement that tax is calculated at checkout by your own jurisdiction. The docs do not reconcile the two, so budget for either. Payment options reportedly vary by region, though no blocked country is named. Verify you can pay before designing around direct access. ## Model IDs Keep Dying Fourteen model IDs have been retired or scheduled for retirement since November 2025, and there is no stable alias to hide behind. For anything with a maintenance window longer than a quarter, that is a bigger operational risk than the price. The [model list](https://platform.kimi.ai/docs/models) records the pattern. `kimi-thinking-preview` was discontinued in November 2025. `kimi-latest` went in January 2026. The entire kimi-k2 series was discontinued on May 25, 2026. `kimi-k2.5` and the whole `moonshot-v1` family, including three `-vision-preview` variants, come off on August 31, 2026. Note what `kimi-latest` was. It was the floating alias, the one string you would pin precisely so you never had to think about this. Moonshot killed it, which means the documented advice is now to pin an explicit version like `kimi-k3` and accept that the explicit version will itself be retired on a timeline you do not control. The practical consequence shows up in the tutorials. Every model string a 2025 guide could have used, `kimi-latest`, the k2 series, `kimi-thinking-preview`, is now dead. Those guides still read as authoritative, because nothing about a dead model ID looks stale until you call it. If you are copying setup code from a blog post, check the model string against the live model list first. For anyone budgeting: assume a migration every two quarters, pin explicit IDs, and keep the model name in configuration rather than scattered through your codebase. ## What Kimi's Pricing Says About AI Visibility The five-engine test from the top of this piece is worth reopening, because the interesting part is why the numbers diverged. Only one of those three tables matched Moonshot's published rate. The two that avoided a wrong number did it by declining to produce a table. One had grounded its search on Moonshot's developer forum instead of a pricing aggregator. Confidence and accuracy ran in opposite directions. Five queries is an anecdote, not a benchmark. The direction is what matters, and the reason for it shows up in the citation data. Sampling AI-engine citations for "kimi api" through DataForSEO's LLM-mentions data on July 18, 2026, `huggingface.co` (65 mentions) and `openrouter.ai` (63) each outranked `platform.kimi.ai` (57), Moonshot's own documentation. Only one pricing page appeared among the most-cited results, and it was an aggregator running four-month-old numbers. So the engines were not hallucinating in the usual sense. They faithfully reported what their sources said, and their sources were resellers and aggregators carrying stale figures. Moonshot publishes correct prices on a well-structured docs site. The engines largely are not reading it. Publishing accurate information is not the same as being the source engines read, and Moonshot is the example. When third parties out-rank you as a citable source for your own facts, the answer machines quote the third parties, and wrong numbers circulate with your brand attached to them. You cannot fix a citation gap you have not found. This is the failure mode geotoolbox measures: which sources engines pull from when they describe you, and where an aggregator has quietly become canonical for your own numbers. Reachability is the first thing to rule out, because a page the crawlers cannot fetch can never be the cited one, and our [AI readiness tool](https://geotoolbox.ai/tools/ai-readiness) checks that in seconds. ## What the Kimi API Really Costs You The sticker price is the smallest part of this. A working integration costs $1 to start, and $10 in cumulative top-ups lifts the account from Tier 0's evaluation limits to Tier 1's. From there the bill is set by which model you pinned and how much output you generate, because output is the part no cache discount touches. The costs that hurt are the ones that are not on the pricing page. A key pointed at the wrong platform. A retry loop backing off against a terminal billing error. A model string that dies on August 31 and takes your integration with it. Budget for a migration every couple of quarters and pin explicit versions, and Kimi is a genuinely cheap way to run coding and agent work. Treat it as set-and-forget and it will surprise you twice, once on the invoice and once at three in the morning. ## Frequently Asked Questions ### Is the Kimi API free? No. A minimum $1 top-up is required before the API will return anything, and there is no permanent free tier. Spending $5 cumulatively earns a $5 voucher, but vouchers do not count toward rate-limit tier progression. The free tier people are thinking of belongs to the consumer Kimi app, which is a separate product from the API. ### How do I get a Kimi API key? Register on Moonshot's developer platform, top up at least $1, and create a key in the console. The step that matters more than the sign-up is choosing the right platform: keys from the international platform, the China platform, and Kimi Code are not interchangeable, and using one against another's base URL fails with a bare 401. ### Does a Kimi subscription include API credits? Not for the metered API. Moonshot bills the API Open Platform separately from the Kimi app and Kimi Code. Kimi Code subscribers do receive their own keys, but those run against a request quota of roughly 300 to 1,200 calls per rolling five-hour window depending on plan, not a token balance. Community threads sometimes claim subscriptions grant matching API credit; Moonshot's documentation does not support that. ### Which Kimi model should I use? For coding and agent work, kimi-k2.7-code at $0.95 input and $4.00 output is the value pick. Choose kimi-k3 when you specifically need the million-token [context window](https://geotoolbox.ai/glossary/context-window) or its stronger reasoning, and price it accordingly, because it costs roughly three times as much per unit of work and is not on the batch pricing table. Avoid starting new work on the legacy moonshot-v1 family: it has no cache discount and retires on August 31, 2026. Some variants still carry lower list prices, but they come with far smaller contexts and almost no migration runway. ### Is Kimi's API cheaper than the alternatives? On the K2 line, yes, and that was Kimi's whole positioning. K3 changes the answer. At $3.00 input and $15.00 output it is priced as a frontier model, not a budget one, so compare it against the frontier models it now competes with, not the cheap tier Kimi used to occupy. Run the numbers on your own input-to-output ratio, and see our [DeepSeek pricing](https://geotoolbox.ai/blog/deepseek-pricing) breakdown for the closest open-weight comparison. ### Can I pay for the Kimi API with a credit card? Moonshot documents WeChat Pay and Alipay for individual accounts and does not mention card payment. Invoices come from Beijing Moonshot AI Technology Co., Ltd. with 6% Chinese VAT. The documentation says payment options vary by account region without specifying which, so confirm you can pay before building against direct access. Buying through a reseller such as OpenRouter is the usual workaround for card-only buyers. ## Sources - Flagship Model Kimi K3 Pricing - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/pricing/chat-k3` - Coding Model Kimi K2.7 Code Pricing - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/pricing/chat-k27-code` - Recharge and Rate Limiting - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/pricing/limits` - Kimi K3 launch top-up rebate (dedicated page since taken down) - Moonshot AI developer docs, as captured July 18, 2026 - `platform.kimi.ai/docs/pricing/promotion` - BatchJob Pricing - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/pricing/batch` - WebSearch Pricing - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/pricing/tools` - Model List and Deprecation Schedule - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/models` - Kimi K2.5 model page - OpenRouter, accessed July 18, 2026 - `openrouter.ai/moonshotai/kimi-k2.5` - Kimi/Moonshot "Rate Limit" error masks insufficient funds - openclaw issue #43447, accessed July 18, 2026 - `github.com/openclaw/openclaw/issues/43447` - Account and Billing - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/guide/account-and-payments` - Kimi Code membership guide - Moonshot AI help center, accessed July 18, 2026 - `kimi.com/help/kimi-code/membership-guide` - Kimi Code documentation - Moonshot AI, accessed July 18, 2026 - `kimi.com/code/docs/en` - API introduction and rate-limit behavior - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/introduction` --- ## Kimi Pricing in 2026: Plans, Free Tier, and Is It Worth It? > Kimi pricing, current to August 2026: the free Adagio tier, all five plans, how the credit system really works, and whether it beats ChatGPT Plus or Claude Pro. - Canonical: https://geotoolbox.ai/blog/kimi-pricing - Published: 2026-07-18 · Updated: 2026-08-08 Kimi pricing starts at zero, and there are four paid tiers topping out at $199 a month. That much is easy. What trips people up is what you are actually buying, because Kimi does not meter its plans in messages or prompts. It meters them in credits, and which model you are talking to decides whether the meter runs at all. Get that wrong, as most guides currently do, and you either overpay for a tier you do not need or budget around limits that no longer exist. Every Kimi figure below was checked against Kimi's own help center on July 18, 2026, and re-checked on July 20. Competitor prices are current as of the same date. **As of July 20, 2026, all four paid Kimi tiers show "Sold out" and no new subscription can be bought.** Moonshot paused new sign-ups on July 19: demand for K3 outran its forecasts and saturated its GPUs. Existing subscribers are unaffected, and Moonshot says places will reopen in batches, with no date given; its [plan-adjustment notice](https://www.kimi.com/help/kimi-blackboard/plan-adjustment-notice) still listed new plans as unavailable to buy as of August 2026. It has also announced that the offer will split in two, a Kimi Membership and a separate Kimi Code Membership. The prices below are still the published ones, but for now they describe a product you cannot purchase. ## How Much Does Kimi Cost? Kimi costs nothing on the free Adagio plan. The four paid plans are $19, $39, $99, and $199 a month, and every one of them is about 20% cheaper if you pay annually. Moonshot names the tiers after musical tempo markings, which is charming and completely unhelpful for working out which one you need. Here is the full ladder, from [Kimi's published pricing details](https://www.kimi.com/help/membership/membership-pricing):
PlanMonthlyAnnual (per month)Annual totalWho it is for
AdagioFreeFree$0Chat and light testing; no K3
Moderato$19$15$180First tier with K3 access and Agent Swarm
Allegretto$39$31$372First tier with Kimi Claw and the 1M context window
Allegro$99$79$948Volume: 360 credits and 4 concurrent tasks
Vivace$199$159$1,908Maximum quotas and the largest swarm allowance
A Kimi membership and Kimi's developer API are separate products with separate bills. Paying for Moderato does not buy you API tokens, and topping up an API balance does not give you a single membership feature. If you are building on Kimi rather than chatting with it, our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) breakdown covers the per-token rates instead. ## What Each Kimi Plan Actually Includes Here is the part almost nobody publishes: the numeric allowances behind each tier, which buyers otherwise crowdsource from forum threads.
AllowanceAdagio (free)Moderato $19Allegretto $39Allegro $99Vivace $199
Agent credits per month660150360720
Concurrent agent tasks12244
Agent speed priorityNone4x4x4x4x
Agent Swarm (beta)None25 uses50 uses120 uses240 uses
Swarm concurrent subtasksNone2448
Kimi Code creditsNone1x5x15x30x
Kimi ClawNoNoYesYesYes
Professional database calls2002,0005,00012,00024,000
Verified against [Kimi's help center](https://www.kimi.com/help/membership/membership-overview) on July 18, 2026. Kimi labels the credit figures "approximate" and "for reference only", so treat them as the vendor's own estimates rather than contractual quotas. Two rows get misreported constantly. Agent Swarm starts at Moderato, not the higher tiers several guides claim. And swarm subtasks top out at 8 concurrent, not the "300 parallel sub-agents" figure that circulates widely: 300 describes a model capability, not a plan allowance, and repeating it overstates what any subscription grants by nearly forty times. One caveat on that row, and it comes from Kimi itself: its pricing page advertises 8 concurrent subtasks for Allegro where its help center says 4. Both pages are Moonshot's, and they contradict each other on the same day. We follow the help center, which is the more precise, tabular source, but check this one if that ceiling is what you are buying. ## Is Kimi Free? What Adagio Actually Gives You Yes. Kimi has a genuine free tier called Adagio, and it needs no credit card. What it gives you is 6 agent credits a month, one agent task at a time, and 200 professional database calls. There is no Agent Swarm, no Kimi Claw, and no speed priority. It also does not include K3: per Moonshot's own [Kimi Code changelog](https://www.kimi.com/code/docs/en/kimi-code/whats-new.html), "Moderato members and above can use Kimi K3." Kimi does not publish which model Adagio serves instead. What it does publish is that K2.6 is the one model whose conversations cost no credits. So whether the free tier's chat is metered comes down to which model you are actually served, and Kimi leaves that unstated. You will also read, on a page Google's AI Overview currently cites and in answers from at least one major AI assistant, that Kimi's free tier gives you 30 to 50 messages a day. It does not. That figure comes from a March 2026 write-up describing an older version of the product, alongside a regional "Pro" plan at $8 to $19 that appears nowhere in Kimi's current tier list. A daily message cap punishes you for talking to the model a lot; a credit pool does not care how much you chat, only how many heavy agent jobs you run. If you picked a plan based on the message-cap version, you probably picked wrong. ## What a Kimi Credit Is, and What Does Not Spend One A credit is a unit of token consumption, not a unit of work. [Kimi's credit rules](https://www.kimi.com/help/membership/update-rules) state that credits are "consumed based on the number of tokens a task processes," so a long document review and a one-line request are not the same purchase, even though both are one task. Everything on a membership draws from a single shared pool: agent tasks like websites, documents, slides, spreadsheets and deep research, plus Kimi Claw, image generation, and chat with the latest model. Kimi Code is the exception, running on [its own separate pool](https://www.kimi.com/help/membership/membership-overview) that scales sharply with tier, from 1x at Moderato to 30x at Vivace. Then there is the exception that changes the upgrade decision, stated in [Kimi's membership overview](https://www.kimi.com/help/membership/membership-overview) and again in its credit rules: **conversations with the K2.6 model do not consume credits.**
![Diagram showing what spends Kimi membership credits: agent tasks, Kimi Claw, image generation and chat with the latest model all draw from one shared pool, Kimi Code has its own separate pool, and chat with K2.6 consumes no credits.](/blog/kimi-pricing/kimi-credit-pool-what-spends-credits.png)
Agent work and latest-model chat draw down a shared pool. Kimi Code has its own. Only K2.6 chat sits outside the meter.
The scope is the whole buying decision. The exemption names K2.6 specifically. Kimi publishes no equivalent for K3, and its credit documentation lists chat with the latest model inside the shared pool. So the free tier's unmetered conversation depends on it serving K2.6, and paying for K3 access is the point at which your chat starts spending credits too. For the jobs that do spend credits, Kimi offers a rough guide, scoped to free-tier users: a simple slide deck runs about 1% to 2% of your credits, a deep research report 5% to 10%, and a code snippet 0.5% to 2%. Treat those percentages carefully, because they do not sit easily beside the plan table. Kimi's own footnote defines those credit figures as "the equivalent number of tasks for the same feature", so Adagio's 6 reads as roughly six agent tasks. Yet a deep research report at 5% to 10% of a pool implies ten to twenty of them. Kimi labels the counts "approximate" and "for reference only", which is a fair warning. Do not budget from either figure: run one representative task and watch what it takes out of your balance. Credits refresh monthly whether you pay monthly or annually, and they land on your subscription anniversary rather than the first of the calendar month. Kimi's own example: subscribe on December 1 at 3:00 PM and they return on January 1 at 3:00 PM. Nothing rolls over, so there is no banking a quiet month to fund a busy one. Tasks can also hit 5-hour and 7-day concurrency limits shown in the interface. Run out mid-month and the task in flight finishes, but new ones are blocked until the refresh. If a task fails on Kimi's end, the thumbs-down button doubles as a credit refund request. ## Monthly vs Annual, and the Downgrade Trap Annual billing saves you about a fifth, and the rate barely moves between tiers.
PlanMonthly12 months at monthlyAnnual totalYou saveEffective discount
Moderato$19$228$180$4821.1%
Allegretto$39$468$372$9620.5%
Allegro$99$1,188$948$24020.2%
Vivace$199$2,388$1,908$48020.1%
That last column is the useful one, and Kimi does not publish it. Work it out from their own figures and the discount is flat at roughly 20% on every tier. Kimi's marketing leads with "save up to $480 a year," which is accurate for Vivace and easy to read as a better deal at the top than it is, because $480 is simply 20% of a bigger number. Several guides report this as "15% to 25%, varies by tier." It does not really vary. Then there is what happens when you change your mind. Upgrading is painless: immediate, prorated, with the unused portion refunded and a fresh credit allocation that ignores whatever you already burned. Downgrading is not. Per [Kimi's plan change rules](https://www.kimi.com/help/membership/membership-upgrade-downgrade), downgrades cannot be applied mid-cycle at all: you cancel, keep your tier until it expires, then subscribe lower. Which is a real argument for starting one tier below what you think you need. Moving up is instant and refunded; moving down costs you the rest of the term. One caveat if you subscribed through the App Store or Google Play: billing changes happen in that store's settings, not on kimi.com, and the cancel and refund rules are the store's rather than Moonshot's. ## Is Kimi Cheaper Than ChatGPT or Claude? At the entry tier, no. Not meaningfully. Kimi is widely described as the budget option, and on the developer API that reputation was earned. On consumer subscriptions it is not, because the whole market has clustered at the same number.
AssistantFree tierCheapest paid tierMain paid tier
KimiAdagio: 6 agent creditsModerato $19Moderato $19
ChatGPTYes, limitedPlus $20Plus $20
ClaudeYes, limitedPro $20Pro $20
GeminiYes, limitedGoogle AI Plus $4.99Google AI Pro $19.99
GrokYes, limitedSuperGrok Lite $10SuperGrok $30
A dollar. That is the gap between Kimi's main paid tier and [ChatGPT Plus at $20](https://geotoolbox.ai/blog/chatgpt-pricing) or [Claude Pro](https://geotoolbox.ai/blog/claude-pricing) at the same price. If you are switching to Kimi to save money on a monthly subscription, you are not saving money. On the API, where the picture changes, our [Kimi K3 vs Claude](https://geotoolbox.ai/blog/kimi-k3-vs-claude) comparison shows K3 winning on cost per task while trailing on speed. [Gemini's pricing](https://geotoolbox.ai/blog/gemini-pricing) undercuts everyone with a $4.99 entry plan, [Grok's](https://geotoolbox.ai/blog/grok-pricing) main tier sits highest at $30, and [DeepSeek](https://geotoolbox.ai/blog/deepseek-pricing) undercuts the lot by giving its app away entirely. Kimi separates itself elsewhere. Its free tier is genuinely usable rather than a demo, and Moonshot publishes open weights for its K2 and K3 models, with [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3)'s weights released on July 27, 2026 under a custom Kimi K3 License. That is a self-hosting path neither ChatGPT nor Claude offers, if you have the GPU infrastructure for it. The "cheap Chinese AI" framing is also under more strain than most guides admit. K3 launched in July 2026 priced at $3 per million input tokens and $15 output on the API, and Artificial Analysis, reviewing it independently, found it [expensive for its performance class](https://artificialanalysis.ai/models/kimi-k3), against a same-price-tier median nearer $1.75 and $9.00. They also note it is very verbose, which compounds the bill because reasoning tokens are billed as output. That is a frontier price tag from the lab that built its reputation on being the affordable one. DeepSeek and Qwen now sit roughly where Kimi used to, which is the comparison that stings, and we cover that field in [Chinese AI models compared](https://geotoolbox.ai/blog/chinese-ai-models-compared). ## Which Kimi Plan Should You Buy? Every tier above Adagio removes one specific constraint, so name the one blocking you. If you cannot, the honest answer is that you do not need to pay yet. One important caveat while the pause holds: this tells you which tier to aim for when sign-ups reopen, not which one to buy today. **Stay on Adagio** if you mostly chat and do not need K3. K2.6 conversations do not spend credits, so if that is what you are being served, a paid plan buys you nothing you are short of. **Moderato at $19** is the first tier that changes anything structural. You get 60 agent credits instead of 6, two concurrent tasks, 4x speed priority, [Agent Swarm](https://geotoolbox.ai/blog/what-is-kimi-ai) at 25 uses, and access to K3. Worth knowing before you upgrade for the model: K3 chat draws on your credit pool in a way K2.6 chat does not. **Allegretto at $39** adds Kimi Claw and lifts Kimi Code credits from 1x to 5x. Kimi Code's changelog also puts the 1M-token context window here, though the pricing page lists it only on Allegro and Vivace, so check that one before paying if the million tokens is what you are buying. It is also usually the plan people mean when they go looking for a "Kimi Pro" subscription. None of Kimi's five tiers is called Pro. **Allegro at $99** more than doubles your credits, from 150 to 360, and is the first tier to lift concurrency past two. **Vivace at $199** is for maxing everything: 720 credits, 240 swarm uses, 24,000 professional database calls. So is a Kimi subscription worth it? Yes, when a measured limit is blocking a workflow you repeat, and the cheapest tier that removes that limit is the one to buy. No, if you expect real savings over ChatGPT Plus, because $19 against $20 is a one-dollar difference, and the free tier may already be doing everything you need. ## Signing Up, Paying, and Cancelling Two practical blockers sit before checkout. Registration is region-gated: Kimi accepts phone numbers and email addresses only from supported regions, and rejects VoIP numbers and landlines, so sign-up fails at verification if your number is not on the supported list. Google sign-in and the app's QR code are the alternate routes people reach for, though neither is documented as bypassing the regional restriction. Payment rails may also be narrower than you expect. Moonshot's [developer-platform documentation](https://platform.kimi.ai/docs/guide/account-and-payments) covers WeChat Pay and Alipay QR top-ups, though that governs API billing rather than membership checkout, and Kimi does not publish an equivalent list for consumer plans. Confirm your card works at checkout before planning around Kimi. Cancellation is in account settings and takes effect at the end of the billing period, or in the app store if you subscribed there. ## Why So Many Kimi Prices Online Are Wrong Most Kimi pricing guides are describing a product that changed underneath them. Kimi has revised its plan structure, its model lineup, and its metering system inside a single year. The pages Google's AI Overview cites for Kimi's free tier are part of the problem: one of the four is a March 2026 write-up still describing a message-capped free tier that no longer exists. We tested how far that drift has spread. On July 18, 2026 we put three pricing questions to four models through their APIs with web search enabled, GPT-5.6, Sonar Pro, Gemini 3.5 Flash and Claude Opus 4.8, plus the live ChatGPT web interface, then checked every answer against Kimi's help center. These were single runs, and model answers vary between runs. Three of the four described Kimi's limits on a daily or weekly cadence rather than the monthly credit pool Kimi operates. Sonar Pro carried a "$9.99 Kimi Pro" tier into its own comparison table, sourced to a third-party listicle. It is the same phantom plan as the "$8 to $19 Pro" above, and no such tier exists. Three repeated the "300 parallel sub-agents" figure against a real ceiling of 8. One, GPT-5.6, did surface the credit exemption and cited Kimi's own help pages for it, though it framed the exemption as covering "its latest models" rather than K2.6 specifically, which is the same mis-scoping this article had to correct. Claude Opus surfaced it too, and was the only one to frame it as a correction to older write-ups, though it got there by way of a third-party site rather than Kimi's documentation. The live ChatGPT surface was precise on the developer API, citing Moonshot's own docs, and named none of the four paid consumer tiers, collapsing the ladder into "starting at $19 per month" and leaning on a tech news review. The gap is specific to the consumer membership, not a general sourcing failure. The models are not broken. They are reading whatever they can reach, and here that is mostly aggregator pages carrying last quarter's numbers. It is the same dynamic we track when we measure [AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) for a brand. Check the date on any Kimi price, and check the number against Kimi's help center rather than a listicle. If a page still talks about daily message caps, it is describing a product Kimi no longer sells. One date still ages part of this page: the K2.5 and moonshot-v1 models retire August 31, 2026. (A summer API top-up rebate has since lapsed, and Kimi K3's open weights, once pending, went public on July 27, 2026 under a custom Kimi K3 License.) The K2.6 credit exemption in particular is a vendor policy about a model that is now one generation behind, so re-check it before relying on it. ## What This Has to Do with Your Own AI Visibility Kimi published its prices clearly, on its own site, in a help center any crawler can read. Months later, several of the most-used AI systems still describe those prices incorrectly, invent a tier, or lean on a news site over the source. The vendor did nothing wrong. It got out-published by aggregators, and the models learned the aggregators. That is the same mechanism that decides how AI answers describe your business. Engines do not verify; they synthesize whatever they can fetch and whatever they absorbed in training. If the accurate version of your pricing is thin on the open web while a stale listing is everywhere, the stale version wins. What decides it is whether a model can reach your page, parse it, and read it without hitting contradictions between sources. That is what [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) is really about, and at geotoolbox we start with reachability rather than wording, because a page an engine cannot fetch will not be fixed by better copy. If you have never checked whether the AI crawlers can read your site at all, start there. The free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) we built at geotoolbox tells you what they can and cannot fetch, before an engine answers a question about you with somebody else's version of the facts. ## Frequently Asked Questions ### Is Kimi AI free? Yes. Kimi's free plan is Adagio, and it needs no credit card: 6 agent credits a month, one agent task at a time, and 200 professional database calls, with no Agent Swarm or Kimi Claw. It does not include K3, and Kimi does not publish which model it serves instead. K2.6 conversations consume no credits, so if that is the model you get, the free tier goes further for chat than for agent work. ### How much does a Kimi subscription cost? Four paid tiers: Moderato at $19 a month, Allegretto at $39, Allegro at $99, and Vivace at $199. Annual billing brings those to $15, $31, $79, and $159 a month, billed as $180, $372, $948, and $1,908 a year. The discount works out at roughly 20% on every tier. ### Is a Kimi subscription worth it? It is worth it when a specific limit blocks something you do repeatedly, and the right plan is the cheapest one that removes that limit. It is not worth buying as a cheaper ChatGPT Plus or Claude Pro, because at $19 against $20 the saving is a dollar. Test on Adagio first. ### Does Kimi have a usage limit? Yes, but not a message limit. Plans are metered by a monthly credit pool that agent tasks, Kimi Claw, image generation and chat with the latest model all draw from, while Kimi Code has a separate pool. Credits refresh on your subscription anniversary and do not roll over. Tasks can also hit 5-hour and 7-day concurrency limits shown in the interface. ### Which Kimi plan gives me K3? Moderato at $19 a month is the first tier with [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3) access, per Moonshot's Kimi Code changelog, which also places the 1M-token context window at Allegretto and above. Kimi's pricing page lists that window only on Allegro and Vivace, so verify it if that is your criterion. Adagio does not include K3, and Kimi does not publish which model it serves instead. Note that K3 chat draws on your credit pool, whereas K2.6 chat does not. ### Does a Kimi subscription include API credits? No. The membership and the developer API are separate products with separate billing. Paying for Moderato does not give you API tokens, and adding funds to an API balance does not grant membership features. If you are building on Kimi rather than using the app, see our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) guide instead. ## Sources - Pricing details (tier prices, annual rates, per-tier allowances) - Moonshot AI, Kimi Help Center, accessed July 18, 2026 - `kimi.com/help/membership/membership-pricing` - Membership plans overview (shared credit pool, K2.6 credit exemption) - Moonshot AI, Kimi Help Center, accessed July 18, 2026 - `kimi.com/help/membership/membership-overview` - Credit update and usage rules (token metering, refresh, rollover, concurrency limits) - Moonshot AI, Kimi Help Center, accessed July 18, 2026 - `kimi.com/help/membership/update-rules` - Membership credit system update (what the shared pool covers, including chat with the latest model) - Moonshot AI, accessed July 18, 2026 - `kimi.com/membership-credits` - Public pricing page and subscription status ("Sold out" markers, announced Kimi / Kimi Code split) - Moonshot AI, observed July 20, 2026 - `kimi.com/membership/pricing` - Moonshot pauses new Kimi subscriptions as K3 demand outstrips GPU capacity - Reuters, July 20, 2026 - Plan changes (upgrade proration, downgrade restrictions) - Moonshot AI, Kimi Help Center, accessed July 18, 2026 - `kimi.com/help/membership/membership-upgrade-downgrade` - Kimi Code changelog, K3 membership availability and the 1M context window - Moonshot AI, accessed July 18, 2026 - `kimi.com/code/docs/en/kimi-code/whats-new.html` - Account and payments, developer platform top-up methods - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/guide/account-and-payments` - Kimi K3 model analysis (independent pricing and cost-efficiency review) - Artificial Analysis, accessed July 18, 2026 - `artificialanalysis.ai/models/kimi-k3` - Flagship Model Kimi K3 Pricing - Moonshot AI developer docs, accessed July 18, 2026 - `platform.kimi.ai/docs/pricing/chat-k3` - Kimi (chatbot), tier names and K3 launch details - Wikipedia, accessed July 18, 2026 - `en.wikipedia.org/wiki/Kimi_(chatbot)` - ChatGPT, Claude, Gemini and Grok subscription prices are stated from each vendor's own published plans and cross-checked against our maintained pricing guides for those products, linked inline, July 2026 --- ## Inkling AI Model: Thinking Machines Lab's First Model, Explained > The Inkling AI model is Thinking Machines Lab's first release: a 975B open-weights multimodal AI. What it is, the specs, the real benchmarks, and who it's for. - Canonical: https://geotoolbox.ai/blog/inkling-ai - Published: 2026-07-17 · Updated: 2026-08-08 The Inkling AI model is the first model Thinking Machines Lab has shipped, and it landed on July 15, 2026. If you have seen the headlines about Mira Murati's secretive lab finally releasing a model and want the plain version of what Inkling actually is, what it can do, and whether it matters to you, this is it, current as of August 2026. Most of the coverage so far is either a spec dump or a launch-day hot take. We will cover the parts they bury: the honest read on how good it really is, the difference between open weights and open source, whether you can run a 975-billion-parameter model at all, and what a new open model means for whether AI tools mention your brand. ## What Is the Inkling AI Model? **Inkling is Thinking Machines Lab's first model: a 975-billion-parameter open-weights [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) model that reasons over text, images, and audio, released on July 15, 2026.** Only 41 billion of those parameters are active for any given token, which is the trick that keeps a model this large affordable to run. The most important thing to understand up front is what Inkling is not. It is not a finished chatbot you sign into, the way you use ChatGPT or Gemini. It is a base model, released as open weights, meant to be downloaded and fine-tuned into something of your own. Thinking Machines is blunt about this: Inkling is a starting point, not a product. That framing echoes the split we describe in our explainer on [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek), another open-weight model that people confuse with a consumer app. There is the model, which is open, and there is the service around it, which is how the company hopes to make money. With Inkling, that service is a fine-tuning platform called Tinker. Keep that distinction in mind and most of the questions about Inkling answer themselves. ## Who Makes Inkling? Thinking Machines Lab and Mira Murati Inkling comes from Thinking Machines Lab, the startup founded in February 2025 by [Mira Murati](https://www.wired.com/story/thinking-machines-lab-releases-its-first-model-inkling/), OpenAI's former chief technology officer and briefly its CEO. She did not leave alone. The founding team includes John Schulman, an OpenAI cofounder who was central to ChatGPT, and Lilian Weng, a former OpenAI VP who led safety and robotics work. It is, in short, a lab built by the people who built a lot of OpenAI. The money followed the names. Thinking Machines raised what was reported as the largest seed round in history, valuing the company at roughly $12 billion before it had shipped a single model. There have also been reports, which the company has not confirmed, of a much larger round near $50 billion that stalled earlier in the year, so treat the eye-watering numbers as reported rather than settled. What the lab had shipped before Inkling was Tinker, its fine-tuning tool, plus some research and a voice-interaction demo. Inkling is the first actual model, arriving almost a year and a half after the company was founded, and the lab has made a point of moving quickly. The culture is deliberately anti-star, favoring team continuity over individual celebrity, which reads as a direct reaction to the environment its founders left. One question the launch does not answer: how a company that gives its flagship away for free, while third parties are free to host the same open weights, eventually pays for the roughly one gigawatt of next-generation Nvidia systems it has committed to. Tinker is the bet, and it is still early. ## Inkling's Specs at a Glance Underneath the positioning, Inkling is a large and modern model. Here is what the [official model card](https://thinkingmachines.ai/model-card/inkling/) and [launch announcement](https://thinkingmachines.ai/news/introducing-inkling/) lay out.
SpecInkling
Parameters975B total / 41B active per token (sparse mixture-of-experts)
Architecture66-layer decoder-only transformer; 256 routed experts (6 active per token) plus 2 shared experts
AttentionInterleaved sliding-window and global layers (roughly 5:1)
Context windowUp to 1M tokens (64K and 256K options on Tinker)
ModalitiesText, image, and audio in; text out only
Training data45 trillion tokens of text, images, audio, and video
OptimizerMuon for large matrix weights, Adam for the rest
NumericsBF16, MXFP8, and NVFP4 support
LicenseApache 2.0 (open weights on Hugging Face)
A few of these matter more than the rest. The mixture-of-experts design is why a 975B model is even practical: it holds a huge number of parameters but only fires a small slice for each [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), so it costs far less to run than a dense 975B model would, even if attention, routing, and long context mean the real bill still lands above a true 41B model. The 1-million-token [context window](https://geotoolbox.ai/glossary/context-window) puts it in the same tier as the largest current models. And it is natively multimodal on input, meaning it was trained from the start to read images and audio, not bolted onto a text model afterward. One caveat worth stating plainly: it only outputs text. It reads images and audio, it does not generate them. ## How Good Is Inkling? Benchmarks and the Caveats Thinking Machines published a full benchmark sheet, run at its highest thinking effort setting. The headline numbers are strong.
BenchmarkInkling scoreWhat it measures
AIME 202697.1%Competition math
GPQA Diamond87.2%Graduate-level science reasoning
SWE-bench Verified77.6%Real-world software fixes
Humanity's Last Exam29.7% (46.0% with tools)Hard expert reasoning
Terminal Bench 2.163.8%Agentic command-line tasks
VoiceBench91.4%Spoken-audio understanding
MMMU Pro73.5%Multimodal reasoning
This is where the launch coverage tends to go quiet. Thinking Machines says outright that Inkling is **not the strongest overall model available today, open or closed**. That is not modesty for its own sake. The frontier closed models from OpenAI, Anthropic, and Google still lead most leaderboards, and on the independent Artificial Analysis Intelligence Index, Inkling scored 41, which made it the top US open-weights model but still left it behind Chinese open leaders like GLM-5.2. Strong for an open model, not a category winner. One more thing about that benchmark sheet: the numbers are reported at the top thinking-effort setting, so they are a best case. Dial the effort down for speed or cost and the scores come down with it. A couple more asterisks. First, several of these scores were produced on the company's own internal test harnesses rather than fully independent ones, so they will need outside reproduction before anyone should treat them as settled. Second, the impressive efficiency claims, like matching Nvidia's Nemotron 3 Ultra on an agentic terminal benchmark with about a third of the tokens, or a fine-tuned Inkling scoring 84.7% on a Bridgewater financial-reasoning test at a fourteenth of the cost, come from vendor or private evaluations. The results may well hold up. They are just not independently verified yet, and it is worth knowing which numbers are which. The first outside checks have started to land. [ARC Prize](https://arcprize.org/results/thinky-inkling), which runs its own evaluation rather than taking a lab's word for it, confirmed on July 17 that Inkling is the highest-scoring open-weight model it has tested on ARC-AGI-1 (79.5%) and ARC-AGI-2 (36.5%). [Sebastian Raschka](https://sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html), an independent machine learning researcher, compared Inkling's published benchmark numbers against GLM-5.2's the day after launch and found a mixed picture. Inkling comes out ahead on instruction-following (IFBench, 79.8% versus 73.3%) and factual accuracy (SimpleQA Verified, 43.9% versus 38.1%). It falls behind on expert reasoning without tools (Humanity's Last Exam, 29.7% versus 40.1%), real-world coding fixes (SWE-Bench Pro Public, 54.3% versus 62.1%), and agentic terminal tasks (Terminal Bench 2.1, 63.8% versus 82.7%). That split result looks more like an honest, broad generalist than a clean sweep in either direction. Raschka's own post cautions that some of those rows mix externally reported figures with the lab's internal-harness numbers, so treat it as a useful outside read rather than a fresh reproduction. ARC Prize's result is closer to what full independence looks like: an evaluation run entirely outside Thinking Machines' control. ## Is Inkling Open Source? Open Weights vs Open Source Inkling is being called open source all over the place, and, as usual, the label is only half right. **Inkling is open weights, not open source.** The weights ship under the [Apache 2.0 license](https://huggingface.co/thinkingmachines/Inkling), one of the most permissive there is, so you can download the model, run it, fine-tune it, and ship it commercially with almost no strings. That is genuinely open, and more permissive than some Chinese open models that add attribution clauses. What you cannot do is rebuild it. Open weights means the finished model files are public. Open source, in the strict sense, would also mean releasing the training data and enough of the recipe to reproduce the model from scratch, and Thinking Machines does not publish its 45-trillion-token dataset. Independent commentator [Simon Willison](https://simonwillison.net/2026/Jul/16/inkling/) flagged the same gap from another angle: the model card's training-data section says only that it uses content "in the public domain as well as content that may be subject to intellectual property protection," without naming sources or explaining where that line was drawn. We walk through why that gap matters in our guide to [open weights vs open source](https://geotoolbox.ai/blog/open-weights-vs-open-source), and it applies squarely here: you can use Inkling freely without being able to fully audit how it was made. There is one more note buried in the technical write-ups. To bootstrap Inkling's early post-training, Thinking Machines generated synthetic data using other labs' open models, including Moonshot AI's [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai) K2.5. Distillation like this is common across the industry, but it is a fair part of the picture: part of how Inkling got good was learning from models someone else trained first. ## Can You Actually Run Inkling? Hardware and How to Access It The spec sheets soften this part. "Open weights" does not mean you can run this on your laptop, or even on a single high-end gaming GPU. Inkling is enormous. The full BF16 checkpoint needs at least 2 terabytes of aggregated GPU memory, which the model card puts at 8 Nvidia B300s or 16 H200s. The quantized NVFP4 version drops that to roughly 600GB, which it still lists as 4 B300s or 8 H200s. The download itself runs to nearly 2TB across a hundred-plus files. This is a data-center model. So for almost everyone, "open" means one of these access routes rather than self-hosting.
How to accessWhat you get
TinkerThinking Machines' own platform to fine-tune Inkling on your data and call it via API. Priced per million tokens (a limited-time launch discount lists around $1.87 prefill / $4.68 output at 64K context; the discount is temporary, so check the current Tinker docs). This is the intended path.
Hugging Face weightsDownload the full open weights and run them on your own cluster. Free license, serious hardware bill.
Third-party providersTogether AI, Fireworks, Modal, Databricks, and Baseten host Inkling as an API so you skip the infrastructure.
Quantized / GGUFNVFP4 for Blackwell GPUs, plus a community 1-bit GGUF via Unsloth for teams squeezing it onto smaller setups, at some cost to quality.
The value in owning the weights is control. Because you can [fine-tune](https://geotoolbox.ai/glossary/fine-tuning) Inkling on your own proprietary data and, on Tinker, keep that data yours rather than folding it into someone else's foundation model, a company can build a specialized model without renting it from a closed lab. That is the pitch, and it is a real one for organizations with valuable internal knowledge. ## Inkling vs DeepSeek, Kimi, and the Open-Model Field Inkling arrives into a field where the best open models have, until now, been Chinese. DeepSeek, Kimi, and Zhipu's GLM set the open-weight bar, and Moonshot raised it further a day later with [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3), a 2.8-trillion-parameter model that scored well above Inkling on the same independent index, and released those weights publicly on July 27. Thinking Machines is pitching Inkling as a serious US entry that claims comparable performance. Here is how it sits against the field.
ModelMakerOpen weights?Strongest at
InklingThinking Machines (US)Yes (Apache 2.0)Multimodal input, 1M context, a customizable base to fine-tune
DeepSeekHigh-Flyer (China)Yes (MIT)Reasoning, coding, very low cost
Kimi (K2 line)Moonshot (China)Yes (modified MIT)Agentic multi-step work, long context
Llama / GLMMeta / ZhipuYes (varies)Broad ecosystems, tooling, adoption
What actually sets Inkling apart is not raw benchmark supremacy. It is the combination of native multimodality, the controllable thinking effort that lets you trade tokens for speed, and the fact that it is a broad, balanced base that a well-resourced team can adapt. If you want the current landscape of the open models it is competing with, our roundup of [how the major Chinese models compare](https://geotoolbox.ai/blog/chinese-ai-models-compared) puts them side by side. Inkling's real claim is geographic and strategic as much as technical: a frontier-scale open model from a US lab, for teams that want an alternative to both the closed American models and the Chinese open ones. ## "Resistance to Censorship": What It Actually Means Some of the launch coverage led with "resistance to censorship," which sounds more dramatic than it is. Thinking Machines frames it under what it calls epistemics: the model should be well-calibrated, follow instructions, and answer politically sensitive questions rather than reflexively refusing them or inheriting a developer's standardized opinions. In plain terms, it is built to be less preachy and less arbitrarily restrictive on contested topics. What it is not is a jailbroken model. Inkling still has hard refusals for genuinely dangerous requests, weapons, cyber abuse, and the like, and it posts strong safety-benchmark scores. The model card is candid about the trade-off, admitting an occasional tendency to comply with role-play or indirectly framed prompts on harmful topics, and it recommends the usual input and output filtering for high-stakes use. One open question worth flagging: because the model is meant to be fine-tuned, how much third-party fine-tuning weakens those built-in safety controls is, by the lab's own admission, still unresolved. ## Should You Use Inkling? The plain answer: Inkling makes sense if you have proprietary domain data worth fine-tuning on, you want to own your model rather than rent one from a closed lab, and you have the hardware to back it up. That means Tinker for most teams, or your own cluster if you need the full million-token context Tinker's shorter windows do not cover. For a bank, a hospital system, or an engineering org with deep internal knowledge and the talent to fine-tune, that is a genuinely appealing package. For everyone else, it is not the tool. If you just want a capable assistant that works out of the box, you do not have a fine-tuning team, or you do not have serious hardware, ChatGPT, Claude, and Gemini remain the better call, and they are more polished besides. Pretending Inkling is something it is not is how people end up disappointed by a model that was never trying to win their use case. ## What Inkling Means for Your AI Visibility While researching this piece we ran a small test. We asked several AI engines "what is Inkling by Thinking Machines Lab." ChatGPT and Perplexity, both connected to live web search, answered accurately and cited the announcement. Google's Gemini, asked without web search, replied that it had "no confirmed information" about the model, and then confidently described a different company by the same name. Just days after a headline launch, a frontier model running on training data alone was not just blank, it was wrong. That is not a knock on Gemini. Without live retrieval, even a frontier model is blind to anything after its training cutoff, and a brand-new launch sits squarely in that blind spot.
![A test of three AI engines asked 'what is Inkling' two days after launch: ChatGPT and Perplexity with web search answered correctly, while Gemini without web search did not know the model and named the wrong company.](/blog/inkling-ai/inkling-ai-engine-knowledge-test.png)
Days after launch, only the web-connected engines knew Inkling existed.
That gap is the whole game in miniature. AI answers are only as current and accurate as the sources a model can reach at the moment it answers. Every user-facing engine, including the products teams will build on open bases like Inkling, is another place a customer might ask "what is the best tool for X" or "is [your company] any good" and act on the reply. In our experience at geotoolbox, the businesses that show up correctly in those answers are rarely the ones with the prettiest homepage. They are the ones a model can find, fetch, and parse without tripping over contradictions, which is the [foundation of generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization). The open-model wave does not change that playbook so much as widen the field. There are simply more engines that can mention, or mangle, what you have built, and the durable move is the same as it has always been: make sure they can read you clearly, then check that they do. Our guides on [what GEO is](https://geotoolbox.ai/blog/what-is-geo) and [tracking your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) go deeper on both. The first step is the cheapest: run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether the AI crawlers can even reach and parse your site before the next model launches and the question gets asked again. ## Frequently Asked Questions ### Is Inkling free? The weights are free to download and use under the Apache 2.0 license, but running them is not: the full model needs a cluster with at least 2TB of GPU memory. If you use it through Thinking Machines' Tinker platform or a third-party provider instead, you pay per token: a limited-time launch discount puts it near $1.87 per million input tokens at the 64K tier, but that discount is temporary, so check current pricing. So the license is free; the compute is not. ### Is Inkling open source? Not in the strict sense. Inkling is open weights: the finished model files are public under Apache 2.0, so you can run, fine-tune, and ship it freely. But Thinking Machines does not release its training data or full recipe, so you cannot reproduce or fully audit how it was built. "Open weights" is the accurate term. ### Can you run Inkling on your own computer? No. Inkling is a 975-billion-parameter model whose full checkpoint needs roughly 2TB of aggregated GPU memory, meaning 8 or more data-center GPUs. A quantized version cuts that to around 600GB, and there is an experimental 1-bit build, but it is still far beyond a single consumer machine. Most people will access it through Tinker or a hosting provider. ### Is Inkling better than ChatGPT or Claude? On raw capability, no. Thinking Machines itself says Inkling is not built to top the leaderboards, and closed models from OpenAI, Anthropic, and Google still lead most benchmarks. Inkling's advantage is that it is open and customizable: you can fine-tune it on your own data and control it, which the closed models do not allow. ### Who created Inkling? Thinking Machines Lab, the startup founded in February 2025 by former OpenAI CTO Mira Murati, along with OpenAI cofounder John Schulman and former OpenAI safety VP Lilian Weng. Inkling, released July 15, 2026, is the lab's first model. ### What is Inkling-Small? Inkling-Small is a preview of a lighter model, 276 billion total parameters with 12 billion active, that Thinking Machines says matches or beats the larger Inkling on several benchmarks thanks to an improved training recipe. As of mid-July 2026 it is only previewed, not fully released. "Small" is relative, though: at 276 billion parameters it is still a data-center-scale model, not something you can run on a single machine. ## Sources - Introducing Inkling (official announcement) - Thinking Machines Lab, July 2026 - `thinkingmachines.ai/news/introducing-inkling` - Inkling model card (official) - Thinking Machines Lab - `thinkingmachines.ai/model-card/inkling` - Thinking Machines amps up its bet against one-size-fits-all AI with Inkling - TechCrunch, July 2026 - `techcrunch.com/2026/07/15/thinking-machines-amps-up-its-bet-against-one-size-fits-all-ai-with-its-first-open-model-inkling` - Thinking Machines Lab Drops Its First Model - WIRED, July 2026 - `wired.com/story/thinking-machines-lab-releases-its-first-model-inkling` - Thinking Machines Lab Unveils Inkling, Its First Open-Weights Multimodal AI Model - Unite.AI, July 2026 - `unite.ai/thinking-machines-lab-unveils-inkling-its-first-open-weights-multimodal-ai-model` - thinkingmachines/Inkling model weights - Hugging Face - `huggingface.co/thinkingmachines/Inkling` - Inkling - ARC-AGI Results - ARC Prize, July 2026 - `arcprize.org/results/thinky-inkling` - Inkling Architecture and Benchmark Notes - Sebastian Raschka, July 2026 - `sebastianraschka.com/blog/2026/inkling-architecture-benchmark-notes.html` - Inkling: Our open-weights model (commentary) - Simon Willison, July 2026 - `simonwillison.net/2026/Jul/16/inkling` --- ## What Is AI Visibility? Definition, Metrics, and Drivers (2026) > AI visibility is how often AI engines like ChatGPT and Google AI Overviews mention, cite, or recommend your brand, plus how to measure and move it. - Canonical: https://geotoolbox.ai/blog/what-is-ai-visibility - Published: 2026-07-17 · Updated: 2026-07-22 Your buyers are asking ChatGPT, Gemini, and Perplexity what to use, and those engines answer with specific brand names. AI visibility is the measure of whether yours is one of them. The term is younger than the problem it describes, so definitions drift, metrics get invented weekly, and the vocabulary around it fights itself. The stable core underneath is smaller than the vocabulary suggests, and it fits on one page. ## What Is AI Visibility? **AI visibility** is how often, and how favorably, AI engines like ChatGPT, Google AI Overviews, Gemini, Perplexity, and [Meta AI](https://geotoolbox.ai/blog/what-is-meta-ai) surface your brand in their answers. It is the AI-era equivalent of search rankings: instead of where you sit in a list of links, it measures whether the AI mentions, cites, or recommends you at all. That definition hides a useful distinction. A **mention** is the AI naming your brand in an answer, with no link attached. A [citation](https://geotoolbox.ai/glossary/ai-citation) is the AI using one of your pages as a source, usually with a clickable reference. The two behave differently: mentions build the recommendation itself, citations carry whatever referral traffic AI search sends. What actually earns them is its own discipline, covered in our guide to [how to get cited by AI](https://geotoolbox.ai/blog/how-to-get-cited-by-ai). So visibility is not a yes-or-no state. In practice your brand sits somewhere on a spectrum: cited and linked, mentioned without a link, or absent. An unlinked [brand mention](https://geotoolbox.ai/glossary/brand-mention) in a ChatGPT answer sends you zero measurable traffic, and it may still be the reason a buyer types your name into Google an hour later. The term travels under aliases. LLM visibility, AI search visibility, and AI brand visibility all describe the same thing: presence inside the answer layer that now sits between your content and a growing share of your buyers. Whatever you call it, the engines are already answering questions about your category. The only open question is whether you appear in those answers. ## AI Visibility vs Traditional SEO Traditional SEO competes for a position in a ranked list. AI visibility competes for a place inside a synthesized answer. That single difference changes the metric, the unit of competition, and the failure mode.
Traditional SEOAI Visibility
What you winA position in a list of linksA mention, citation, or recommendation inside the answer
Unit of measurementRanking position, clicks, CTRMention rate, citation rate, share of voice, sentiment
How stable it isRankings shift over weeksThe same prompt can produce a different answer an hour later
How you check itRank trackers, Search ConsoleRepeated prompt sampling per engine, plus emerging first-party reports
Failure modePage two obscurityAbsence: the answer simply never names you
The two used to be almost the same discipline. Ahrefs has tracked the share of Google AI Overview citations that come from top-10 ranking pages: 76% in mid-2025, down to 38% by early 2026 in [its 863,000-SERP update](https://ahrefs.com/blog/ai-overview-citations-top-10/). Part of that drop is Ahrefs parsing more citations than it used to, and Ahrefs' conclusion held anyway: ranking for the exact query no longer guarantees a seat in the answer, and most of what AI Overviews cite now lives outside the top 10. That decoupling explains the contradictory numbers floating around this topic. Whether ranking correlates with AI citations depends on when, and how, the study measured. Across the snapshots available, the direction is consistent: the engines increasingly pick sources by [how they retrieve and synthesize](https://geotoolbox.ai/blog/how-does-ai-search-work), not by who ranks first. SEO is still the on-ramp. Reachable, well-structured, well-ranked content remains the raw material engines pull from. It just stopped being the whole game. One mechanical detail worth knowing: your brand enters an AI answer through two doors. **Training data** bakes in whatever the web said about you months ago, which favors established brands with long histories. **Live retrieval** pulls current pages at answer time, which is the door a newer brand can influence this quarter. Every layer in the drivers section below works on the second door. Whether and when it reaches the first depends on the model makers' training runs. ## Why AI Visibility Matters in 2026 The behavior shift is documented, and it is not subtle. When [Pew Research Center analyzed 68,879 real Google searches](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/), users clicked a traditional result on just 8% of visits when an AI summary appeared, versus 15% without one. Links inside the AI summary itself were clicked on roughly 1% of visits. Read that last number again. Even when the AI cites you, almost nobody clicks through; in Pew's data the answer surface is nearly zero-click. The value of appearing in the answer is mostly not the click. It is being the brand the answer names when a buyer asks what to use, what to trust, or what to compare. AI Overviews now show on roughly half of US searches by tracker estimates, and the standalone assistants keep growing on top of that. The full engine-by-engine numbers live in [State of AI Search 2026](https://geotoolbox.ai/blog/state-of-ai-search-2026). The other half of the picture: AI referral traffic is still small, somewhere between a rounding error on broad-web referrals and low single digits of B2B inbound, depending on whose panel you trust. What makes it interesting is intent. Visitors who arrive from an AI answer already got the summary and clicked anyway. If your buyers are not asking AI engines about your category yet, AI visibility is not your bottleneck. Run a handful of real buyer prompts through ChatGPT and Gemini before you spend a quarter optimizing for them. Measure first, then decide. ## What Drives AI Visibility? Three layers, in order. Most advice on this topic starts at layer two or three and skips the one that decides whether your own pages can show up at all.
![Diagram of the three layers that drive AI visibility, in order: crawler reachability, machine-readable structure, and third-party corroboration](/blog/what-is-ai-visibility/what-drives-ai-visibility-stack.png)
Engines must be able to fetch you before structure or reputation can matter.
**Layer 1: reachability.** Before your own pages can be cited or your current messaging retrieved, the engine's crawler has to fetch your website. A robots.txt rule blocking retrieval crawlers like OAI-SearchBot or PerplexityBot, a WAF or CDN challenge that swallows bot requests, or content that only renders in JavaScript takes your pages out of the citation pool no matter how good they are. Blocking GPTBot, OpenAI's training crawler, is a separate decision: it shapes what future models bake in, not whether today's answers can cite you. Either way, your brand can still surface through third-party coverage, but what the AI says about you is then hostage to what everyone else publishes. It is an easy failure to have without knowing: the website looks fine in a browser, and nothing in a marketing dashboard says otherwise. Our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) lists the user agents worth checking, your server logs show which of them are already hitting you, and our free [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) tests your site against them in about a minute. **Layer 2: machine-readable structure.** Engines extract passages, not pages. Content with a direct answer up front, self-contained sections, real HTML tables, and structured data gets lifted into answers more reliably than clever prose. Freshness helps too: in the datasets that measure it, recently updated pages hold citations better than abandoned ones. The specifics live in our guide to [optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search). **Layer 3: third-party corroboration.** On the evidence so far, the engines weigh what the rest of the web says about you more than what you say about yourself. In the [AirOps 2026 State of AI Search dataset](https://www.airops.com/report/the-2026-state-of-ai-search), about 85% of brand mentions in AI answers came from external domains, and community and user-generated content (UGC) platforms alone accounted for roughly 48% of citations. That is one vendor's dataset, so treat the exact numbers as directional, but independent analyses keep landing in the same place: earned presence, reviews, and [entity clarity](https://geotoolbox.ai/blog/entity-seo) are where the leverage concentrates. The same signals that build [E-E-A-T for AI search](https://geotoolbox.ai/blog/eeat-ai-search) build this layer. ## How Do You Measure AI Visibility? Not with a rank tracker, and not with a single number. AI visibility tracking means sampling answers across platforms and scoring what comes back. The working metric set looks like this:
MetricWhat it answersThe nuance
Mention rateOut of your tracked prompts, what share of answers name your brand at all?The base presence metric
Recommendation rateWhen the answer shortlists options, how often are you recommended rather than merely named?The commercial metric; a mention in a caveat is not a recommendation
Citation rateHow often do your own pages get used as a source?AI citations carry the referral traffic
AI share of voiceOf all brand mentions in your category's answers, what share is yours vs competitors?AI share of voice, explained
Position and prominenceFirst recommendation, mid-list, or a footnote?Weight your scoring toward first mentions
Sentiment and accuracyDoes the AI describe you correctly and favorably?Wrong facts in answers are their own problem to track
The part most coverage skips: **one measurement of any of these is meaningless.** AI answers are probabilistic. In the AirOps dataset above, only 30% of brands stayed visible from one answer to the next, and just 20% persisted across five consecutive runs. The exact numbers are one vendor's, but the variance itself is not in dispute: an April 2026 paper puts the method plainly in its title: [Don't Measure Once](https://arxiv.org/abs/2604.07585). Visibility is a distribution, and you have to sample it: same prompts, per engine, repeated over time. The trend is the signal; any single run is noise. Disclosure: we build one of these trackers, and this is why geotoolbox reports presence per engine over time instead of one blended score. Dashboards inherit that noise, plus one of their own. Many tracking platforms query the engines through APIs, and API answers are not always what a logged-in user sees in the app, where memory, model routing, and search grounding differ; platforms that scrape the interface instead trade that problem for scraping fragility. Marketers have noticed: as one exec told [Digiday](https://digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/), run the same prompts through three tools and you get three different answers. Tools are benchmarkers, not ground truth, ours included. Our walkthrough of [how to track AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) covers a sampling protocol you can run before paying for anything. ## Can You See AI Visibility in Your Own Analytics? Partially, and 2026 is the first year that answer is not a flat no. Referral data catches a slice. Sessions arriving on your website from chatgpt.com, perplexity.ai, or copilot.microsoft.com show up in analytics as referrals, and they are worth segmenting. But they undercount badly: plenty of AI-influenced visits arrive as direct traffic or branded search, because the user read the answer, closed the tab, and looked you up later. The bigger change is official first-party data. In June 2026, Microsoft expanded [its AI visibility reporting in Bing Webmaster Tools](https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare) with query intents, topic clusters, and a citation-share view showing your slice of all citations for the same grounding queries, on top of the Copilot and Bing citation counts it began publishing in February. Google began rolling out [generative-AI performance reports in Search Console](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports) the same month, showing impressions in AI Overviews and AI Mode, initially to a subset of sites and without click data. For the first time, part of your AI visibility is observable from the platform side instead of reconstructed through synthetic prompts. Neither surface covers ChatGPT, Claude, or Perplexity. For those, sampling remains the only window. And if you need a proxy that leadership already trusts, watch branded search volume, the closest thing this channel currently has to attribution. ## AI Visibility, GEO, AEO: Which Word Means What? The vocabulary is messier than the concepts. Here is the map: **AI visibility is the outcome.** It is the measurable state of appearing in AI answers: the metrics in the table above. **[GEO](https://geotoolbox.ai/blog/what-is-geo), AEO, and LLMO are the practice.** [Generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization), answer engine optimization, and large language model optimization are three labels the industry coined for the same job: making your brand and content more likely to be retrieved, cited, and recommended by AI engines. The overlap between them is roughly total; the differences are mostly about who is selling what. We break down the labels in [GEO vs AEO vs SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo). So the sentence that keeps the terms straight: you do GEO to improve your AI visibility, the way you do SEO to improve your rankings. ## How to Improve AI Visibility The full playbook belongs to the guides linked from each layer above; what this page owes you is the order. Reachability first, because nothing downstream works without it. Liftable, answer-first pages second. The third-party layer third, and it is where the data above says most of the leverage lives. Measurement wraps around all of it: check your mention rate before and after each move, or you are just shipping vibes. In our experience, the measurement step is where teams quietly fail. They check ChatGPT once, see themselves mentioned, and declare victory, when five runs of the same prompt would have shown them appearing twice out of five. Build your process around that variance, and once a quarter, run the whole loop as a structured [AI visibility audit](https://geotoolbox.ai/blog/ai-visibility-audit). ## Frequently Asked Questions ### What is a good AI visibility score? There is no standard scale, and any single blended score hides the variance that matters. Track mention rate per engine over repeated runs instead: in the categories that get measured, a handful of brands tend to dominate the answers while everyone else sits near zero, so your competitors' rates are the benchmark that means the most. Our guide to the [AI visibility score](https://geotoolbox.ai/blog/ai-visibility-score) explains what a useful score is built from. ### Can you rank #1 on Google and still be invisible in AI answers? Yes, and it is increasingly common. In Ahrefs' tracking, the share of Google AI Overview citations coming from top-10 pages fell from 76% in mid-2025 to 38% in 2026 (a shift Ahrefs partly attributes to its own improved citation parsing), so ranking and AI citation have visibly decoupled. Ranking well still helps, but it no longer guarantees a place in the answer. ### How often should you check your AI visibility? Weekly sampling is the practical floor, because answers change run to run: in AirOps' 2026 dataset, only about 20% of brands stayed visible across five consecutive runs of the same prompt. Check trends over weeks, not single snapshots on any given day. ### Is AI visibility the same as brand monitoring? No. Brand monitoring tracks mentions of your brand across the web, social, and press. AI visibility tracks whether AI engines mention, cite, or recommend you inside their generated answers, which is a different surface with its own metrics and its own failure modes. ### Do you need a tool to track AI visibility? Not to start. Running 10-20 real buyer prompts across ChatGPT, Gemini, and Perplexity a few times per week in a spreadsheet will tell you where you stand. Tools earn their fee at scale, and they disagree with each other enough that a manual baseline keeps them honest. Our comparison of the [best AI visibility tools](https://geotoolbox.ai/blog/best-ai-visibility-tools) covers when paying makes sense. ### Which AI platforms matter most for visibility? Start with Google's AI surfaces and ChatGPT, then add Perplexity, Gemini, Claude, and Copilot as your tracking budget justifies it. The platforms behave differently: when we sampled citation data for this topic in July 2026 (Google AI surfaces plus ChatGPT's cited sources), Google's AI answers leaned on tool listicles while ChatGPT cited definition and glossary pages. Measure each one on its own terms. ## Where to Go from Here So, what is AI visibility? It is presence in the answer layer: how often the engines your buyers now consult mention, cite, and recommend you, measured per engine, over repeated runs. Start with your own baseline, because every strategy decision downstream depends on it. Our free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) tests the foundation: whether AI engines can actually reach and parse your site. It takes a few minutes rather than a tool subscription. If the foundation checks out, the measurement and improvement guides linked throughout this page are the path from there. We built geotoolbox to run exactly that loop: sample the engines, track the trend, and show you what moved. ## Sources - Google users are less likely to click on links when an AI summary appears in the results - Pew Research Center, July 2025 - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/` - AI Overview citations from top-10 pages study - Ahrefs, March 2026 - `ahrefs.com/blog/ai-overview-citations-top-10/` - Don't Measure Once: Measuring Visibility in AI Search (GEO) - arXiv, Schulte, Bleeker & Kaufmann, April 2026 - `arxiv.org/abs/2604.07585` - New AI Visibility Insights in Bing Webmaster Tools - Microsoft Bing Blog, June 16, 2026 - `blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare` - Marketers question expensive AI visibility tools as inconsistent results fuel skepticism - Digiday, Kimeko McCoy, May 2026 - `digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/` - The 2026 State of AI Search - AirOps - `airops.com/report/the-2026-state-of-ai-search` - Introducing Search Generative AI performance reports in Search Console - Google Search Central Blog, June 3, 2026 - `developers.google.com/search/blog/2026/06/gen-ai-performance-reports` --- ## What Is Kimi K3? Moonshot AI's 2.8T Open Model, Explained > What is Kimi K3? Moonshot AI's open-weight model explained: real specs, which benchmarks to trust, pricing, whether you can run it, and K2 vs DeepSeek. - Canonical: https://geotoolbox.ai/blog/what-is-kimi-k3 - Published: 2026-07-17 · Updated: 2026-08-21 Ask a current AI assistant what Kimi K3 is and it will show you the problem in real time. With web search switched off, Gemini answers "I do not have reliable, verified information about a model called Kimi K3," and Claude says much the same. The model launched on July 16, 2026. The assistants most people rely on have not caught up, and will not for months. So here is the plain version, current as of July 2026: what Kimi K3 actually is, whether the launch-day hype survives contact with the numbers, what it costs, whether you can run it, and how it stacks up against Kimi K2 and DeepSeek. We will flag which claims are Moonshot's own and which have been checked by someone independent, because on a launch-week model that distinction is most of the story. ## What Is Kimi K3? **Kimi K3 is Moonshot AI's 2.8-trillion-parameter open-weight [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) model, launched on July 16, 2026, with a one-million-token context window. It is the largest open-weight model any lab has released, and the weights went public on July 27 under Moonshot's own Kimi K3 License.** Only a small fraction of those parameters fire on any given token, which is the design trick that makes a model this size runnable at all. The more useful thing to understand is what Kimi K3 is not. It is not a finished, sign-in-and-chat product that beats everything else, and it is not the budget option Moonshot built its name on. It is a specialist: strong on coding, agent workflows, and long-context tasks, priced and positioned as a frontier model, and noticeably weaker outside that lane. One developer testing it on launch day put it bluntly, calling it "materially worse than GPT-5.6 Sol and Fable 5 for non-coding use cases." Kimi K3 sits at the top of the same lineup we cover in our [Kimi AI explainer](https://geotoolbox.ai/blog/what-is-kimi-ai), which walks through the K2 family, the open-weights model, and the safety and China questions in more depth. This article is about K3 specifically: the flagship release that pushed Moonshot from "the cheap Chinese open model" into direct frontier competition, and what that shift means once you look past the headline benchmarks. ## Who Makes Kimi K3? Moonshot AI and the "DeepSeek Moment" Kimi K3 comes from Moonshot AI, a Beijing lab founded in 2023 by Yang Zhilin and backed by Alibaba, Tencent, and Meituan. The company has climbed fast: valued near $4 billion at the end of 2025, it raised about $2 billion in May 2026 at a valuation of about $20 billion, and it is reportedly raising again at roughly $30 billion ahead of a Hong Kong listing. That is a steep curve for a lab whose reputation was built almost entirely on giving its models away. The launch landed as a geopolitical event, not just a product release. Reuters framed K3 as the world's largest open-weight AI system and reported that it arrived weeks after the U.S. government abruptly withdrew Anthropic's Fable and Mythos models over security concerns, and that shares in Chinese rivals Zhipu and MiniMax fell sharply on the news, down 27.7% and 16.5% in Hong Kong. Until K3, Meituan's LongCat-2.0 and DeepSeek's V4-Pro had led the field at around 1.6 trillion parameters. Alibaba answered within days by previewing [Qwen3.8-Max](https://geotoolbox.ai/blog/qwen3-8-max), a claimed 2.4-trillion-parameter flagship whose text weights it later opened, on August 12, 2026, under a bespoke license. The reaction split along a familiar line. The loud version, common on launch day, was that the gap between Chinese and U.S. labs has all but closed. The sober version, which we find more defensible, is that K3 narrows it to under three months on the tasks it is strongest at. Either way, neither settles whether the benchmark wins actually hold up. ## Kimi K3 Specs and Architecture Kimi K3 holds 2.8 trillion parameters in total but activates only 16 of 896 experts for any given token, so the compute cost per token stays far below what the full size suggests. That sparse routing runs on a framework Moonshot calls Stable LatentMoE, which lets the model push that aggressive ratio without the training instability that usually comes with it. Moonshot has published the expert counts but not the active-parameter figure in billions, so treat any exact "active parameters" number you see as an estimate. This routing of tokens through a handful of specialist sub-networks is the same mixture-of-experts pattern we break down in [how ChatGPT works](https://geotoolbox.ai/blog/how-does-chatgpt-work). The headline architecture change is Kimi Delta Attention, a hybrid linear-attention design that, in Moonshot's own Kimi Linear research, decodes up to 6.3 times faster at long context. It is paired with Attention Residuals, which pull information across model depth for a claimed efficiency gain at minimal extra cost, and Gated MLA for sharper attention. Roughly three out of every four attention layers use the cheaper linear form, which is what cuts the memory footprint enough to make a 1M window practical rather than theoretical. The practical specs matter more than the internals for most readers. Kimi K3 takes text, images, and video as input and returns text, with a [context window](https://geotoolbox.ai/glossary/context-window) of 1,048,576 tokens and default output up to 131,072. Reasoning is always on, and the `reasoning_effort` control now accepts three levels, `low`, `high`, and `max`, defaulting to `max`; the sampling parameters are locked server-side. At launch only `max` was available, so early write-ups describe a single-setting model. Put simply: K3 always thinks, and unless you dial the effort down it thinks hard on every request, which shows up in both the speed and the bill. ## The Benchmarks, and Which Ones You Can Actually Trust Here is the part most launch coverage skips. On a launch-week model, nearly every eye-catching number comes from the lab that built it, run under conditions the lab chose. That does not make the numbers wrong, but it does mean you should sort them by who measured them before you draw conclusions.
![Table sorting Kimi K3 benchmarks by who measured them: independent results (AA Intelligence Index 57, Frontend Code Arena #1, hallucination rate 51%) versus Moonshot's vendor-reported scores (GPQA Diamond 93.5, Terminal-Bench 88.3, BrowseComp 91.2).](/blog/what-is-kimi-k3/kimi-k3-benchmarks-vendor-vs-independent.png)
On a launch-week model, the split between independently verified and vendor-reported numbers is most of the story.
The results worth leaning on today are the independently checked ones. [Artificial Analysis](https://artificialanalysis.ai/models/kimi-k3) puts K3's overall Intelligence Index at 57. That ranked it fourth of 189 models at launch; as of July 25, 2026 it sits seventh of 190, behind Claude Opus 5, Fable 5 and GPT-5.6 Sol, having been passed by the newer Anthropic and OpenAI configurations rather than by any drop in its own score. Notably it still edges the legacy Opus 4.8, which scores 56. Its long-horizon knowledge-work Elo of 1547 trailed only Fable 5 when measured at launch. Separately, the Frontend Code Arena, a human-preference leaderboard, ranks it first, ahead of Fable 5 and GPT-5.6 Sol, though critics note that leaderboard leans heavily on frontend and 3D-demo tasks, so read it as coding-flavor strength rather than general capability. Read together, those say something specific: K3 is at the frontier on narrow coding and agent tasks, and merely competitive on general intelligence. Everything else in the headline tables is Moonshot's own reporting, and the caveats are real. The company ran different benchmarks through different agent harnesses (its own Kimi Code, Claude Code, or Codex) at maximum thinking effort, so the comparisons are not strictly apples to apples, and the author of one benchmark Moonshot cited publicly objected that the metric can inflate partial-credit scores. There is also a result the vendor page does not headline: on Artificial Analysis's hallucination test, K3's fabrication rate rose to 51% from the previous model's 39%, even as its accuracy improved. It answers more questions correctly and makes up more of the ones it gets wrong.
BenchmarkKimi K3Fable 5GPT-5.6 SolMeasured by
AA Intelligence Index57 (#7 of 190, July 25)higherhigherIndependent (Artificial Analysis)
Frontend Code Arena1,679 (#1)1,6311,618Independent (Arena)
GPQA Diamond93.592.694.1Vendor-reported
Terminal-Bench 2.188.384.688.8Vendor-reported
BrowseComp91.288.090.4Vendor-reported
HLE (general reasoning)43.553.344.5Vendor-reported
Hallucination rate51% (up from 39%)54.9% (higher)n/aIndependent (Artificial Analysis)
Even Moonshot's own table has K3 trailing Fable 5 on general reasoning, as the HLE row shows, so this is a coding and agent specialist rather than an across-the-board leader. Believe the independent numbers, treat the vendor table as a claim awaiting reproduction, and expect independent coding and reasoning benchmarks to fill in over the coming weeks. ## Kimi K3 Pricing: "Open" Does Not Mean Cheap The biggest surprise of the launch was not a benchmark. It was the price. Kimi K3's API costs $3.00 per million input tokens, $0.30 per million on a cache hit, and $15.00 per million output tokens. Moonshot has since published an official USD rate card confirming those figures; we break down the full lineup, the tier system, and the access routes in our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) guide. If you are weighing the consumer app rather than the API, the subscription tiers and the credit system are covered in our [Kimi pricing](https://geotoolbox.ai/blog/kimi-pricing) guide. Against Moonshot's own history, the jump is stark. That output rate is nearly four times what the [K2.7 Code](https://geotoolbox.ai/blog/kimi-api-pricing) model charged, and the input price is more than three times higher. As one widely shared reaction put it, this is "frontier pricing, from the lab whose entire identity was being the cheap one." The era of a Chinese open model automatically being the budget pick is over.
ModelInput / 1MOutput / 1MCache-hit inputNote
Kimi K3$3.00$15.00$0.30Frontier tier; nearly 4x the K2 line's output
Kimi K2.6 / K2.7 Code$0.95$4.00discountedStill open, far cheaper
DeepSeek V4 Pro$0.66 off-peak / $1.32 peak$1.98 off-peak / $3.96 peak~$0.022Frontier-class, a fraction of K3 per token
Whether $15 hurts depends entirely on your workload. Because K3 reasons on every request and can be verbose, a single task often drags a long thinking trace, retried tool calls, and a growing history through the output meter, so heavy output bills are routine rather than rare. On Artificial Analysis's cost-per-task measure K3 runs about $0.95, under half the $2.03 [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5) costs at max effort and close to GPT-5.6 Sol at $1.04, which is reasonable for frontier-grade work. But if your goal was to save money by going open, note that DeepSeek V4 Pro completes a comparable task for a small fraction of that. Our [DeepSeek pricing breakdown](https://geotoolbox.ai/blog/deepseek-pricing) and [Claude pricing guide](https://geotoolbox.ai/blog/claude-pricing) give the fuller comparison. ### The 1M Context Is Smaller Than It Looks, Twice Over The headline context window and the one you can actually reach differ for two separate reasons, and neither is on the spec sheet. The first is a payload ceiling. Moonshot's own issue tracker carries a reproducible `400 total message size 2100954 exceeds limit 2097152` error: a hard 2 MiB cap on total conversation size. In the reporting user's session that worked out to roughly 770,000 tokens, well short of the advertised 1,048,576. So on text-heavy conversations the byte cap bites before the token limit does. The second is a paywall. On the Kimi coding plans, 1M context is a tier feature rather than a model property. K3 is not available on the free Adagio or entry Andante tiers, which stay on K2.7; K3 access starts on Moderato (¥99), and the full 1M-token window is exclusive to the top Allegro (¥699) tier. If you are subscribing specifically for the million-token window, check which tier actually grants it before you buy, because the cheapest plan does not give you K3 at all. ### The Subscription Burn Rate Is the Part That Surprised People Per-token pricing is only half the story, and it is not the half that bit early users hardest. The consistent first-week complaint across Kimi's coding plans is that a single task can consume a startling share of a usage window. One user on the $19 plan ran a task he benchmarks every model against and watched it eat almost his entire five-hour allowance, where the same task on a $20 OpenAI plan finished in minutes and barely registered. Another on the same entry tier hit a loop retrying a Docker step and burned through a five-hour window, which was 20% of his weekly quota, on that one failure. A third, on the $99 Kimi Coding plan, reported quota draining at a pace similar to a $200 Anthropic subscription, and notably he liked the model, rating it above Opus 4.8 on quality. The mechanism is measurable. Artificial Analysis needed 130 million tokens to run K3 through its Intelligence Index against a 63 million average across the field, roughly double. In Simon Willison's test, a single SVG generation returned 16,658 output tokens of which 13,241 were reasoning, costing 25 cents for one image. Several users independently describe the same trace pattern behind it: paragraphs of "wait, actually" as the model backtracks and second-guesses itself on small details. This is where `reasoning_effort` stops being a footnote. At launch "max" was the only accepted value, with no cheaper mode to fall back to; Moonshot has since added `low` and `high`, but `max` is still the default, so every request you do not explicitly turn down pays full reasoning tokens at the $15 output rate. Max effort also makes K3 slow enough to break evaluation harnesses: one public comparison had to raise a five-minute per-task timeout to thirty, and one task still took around nine minutes. If cost or latency matters, set the effort level explicitly rather than leaving it at the default. One more thing worth naming: almost nobody can measure their own burn. The plans report usage as an opaque percentage rather than tokens, which is why some users route their coding subscriptions through a gateway purely to get visibility into what they are actually spending. ## Can You Actually Run Kimi K3? "Open weights" sounds like you can download K3 and run it yourself. The weights are genuinely public now, released on July 27 to Moonshot's Hugging Face repo, and community quantizations for llama.cpp, Ollama, and LM Studio appeared almost immediately. But for almost everyone, downloadable is still not the same as runnable, and the reason is scale. At 2.8 trillion parameters, the repository is about 1.5TB to download in K3's native 4-bit MXFP4 format, and you need roughly 1.4TB of fast accelerator memory just to load the weights before any conversation. Moonshot recommends serving it on a supernode of 64 or more accelerators; community estimates put the floor around 21 H100-class GPUs, or three server nodes of eight 80GB cards each. A single RTX 4090 or a 512GB Mac Studio does not come close. One tester who did get it onto an M1 MacBook, by streaming individual experts from Hugging Face per token, measured it at roughly one minute per token, which captures the gap between "I can technically load it" and "I can use it." This is the tension practitioners keep circling: K3 may be legally open while staying operationally closed to anyone without a data center. What open weights buy you here is not laptop inference. The most-upvoted framing in the community is that the real payoff is provider competition: because anyone can host K3, a market of API providers drives the price down, the way many independent hosts already do for models like GLM. Alongside that you get durability (a version cannot be silently retired out from under you) and data residency for teams that self-host. So who self-hosts? Organizations with real GPU infrastructure and a reason to keep data in-house, and even they mostly run it non-interactively. That is the practitioners' fix for the speed problem: point it at an overnight job like "analyze this codebase for vulnerabilities" rather than an interactive chat. A response an hour later is fine there. For everyone else, the practical paths are Moonshot's own platform, a router like OpenRouter, or the free tier in the Kimi consumer app. If you do want to try self-hosting, our [guide to running Kimi K3 locally](https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally) walks through every route, the real hardware, and what each one costs. Note that API credits are billed separately and are not bundled into any Kimi app subscription, so paying for the app does not hand you API access. Our [open weights vs open source](https://geotoolbox.ai/blog/open-weights-vs-open-source) explainer covers why the distinction matters, and our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking includes a self-host reality check for this class of model. ## Kimi K3 vs Kimi K2 vs DeepSeek: Which Should You Use? The short answer: reach for Kimi K3 only when the job specifically needs what it does best, and keep something cheaper or more reliable for everything else. K3 earns its 3-to-4x premium when you need the long context (bearing in mind the 2 MiB payload cap and the plan tiering above), native vision, or frontier-grade coding and agent performance, and when you can tolerate its speed, which runs a modest 28 to 62 tokens per second with a lot of thinking in between. For most day-to-day work the math favors its own siblings. Kimi K2.6 and K2.7 Code are far cheaper, fully open, and already strong on coding, so unless a task hits K3's specific strengths, the older models do it for a fraction of the cost. If you are optimizing purely for price per token, [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek) V4 is cheaper still. And where the cost of a wrong answer is high, Claude Fable 5 and Opus keep the edge on careful reasoning and verified coding. A fuller side-by-side of the open Chinese models lives in our [Chinese AI models comparison](https://geotoolbox.ai/blog/chinese-ai-models-compared).
ModelBest forOpen weights?Rough costWatch for
Kimi K31M context, vision, frontier coding and agentsYes (public, custom license)High ($15 output)Slow, verbose, needs a GPU cluster to self-host
Kimi K2.6 / K2.7 CodeEveryday coding at low costYesLowNot frontier-level on the hardest tasks
DeepSeek V4Cheapest reasoning and coding per tokenYesLowestSame China data questions
Claude Fable 5 / OpusHigh-stakes reasoning and verified codingNoPremiumClosed; you rent, not own
The mature move is to pilot, not switch. Run K3 on one real, measurable task alongside your current model, look at accepted results and how much supervision each needed, and keep whichever leaves less total friction. On launch-day evidence, K3 deserves that pilot for coding, agents, and long-context work. It does not yet justify replacing a model you trust for everything. ## What Kimi K3 Means for Your AI Visibility Come back to where this started. A day after launch, the assistants most people use could not describe Kimi K3 because their training predates it, and they will stay behind for months. That lag is not a Kimi quirk. It is how every model treats anything new, including your business. If a brand-new, heavily covered AI model is invisible to deployed assistants, then a product launch, a rebrand, or a corrected fact about your company is invisible the same way until training and the live web catch up. That gap is exactly what [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) work addresses. And because K3's weights are public, the model gets fine-tuned and embedded into a long tail of downstream tools you will never see individually, each answering questions about your market from whatever it can find about you. That makes two things the actual levers, and neither is the model. The first is reachability: every one of these systems and the crawlers feeding them has to be able to fetch your site, or you are absent from the live layer that updates faster than training does. The second is consistency, the [core of getting cited by AI](https://geotoolbox.ai/blog/what-is-geo): the businesses described correctly are the ones whose facts line up across the sources a model reads. In our experience at geotoolbox, the companies that surface well in AI answers are rarely the ones with the prettiest homepage; they are the ones a model can find, parse, and trust without tripping over contradictions. You cannot control what Kimi K3 or the next open model learns about you. You can control whether it can reach you at all. Run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether the AI crawlers can fetch and parse your site, and fix the gaps before the next launch makes the question urgent again. ## Frequently Asked Questions ### Is Kimi K3 Chinese? Yes. Kimi K3 is built by Moonshot AI, a Beijing lab founded in 2023 and backed by Alibaba, Tencent, and Meituan. Because it is a Chinese company, the same data-jurisdiction questions that apply to any China-hosted service apply here, which is one reason the open weights matter: teams that need to keep data in their own jurisdiction can self-host rather than send prompts to Moonshot's servers. As with K2, no independent safety evaluation of K3 has landed yet, so treat its alignment and refusal behavior as unverified for now. ### Is Kimi K3 open source? Not quite. K3 is open weights, not open source. Moonshot released the model files on July 27, 2026 under its own Kimi K3 License, which permits commercial use but is not an OSI-approved open-source license: it adds a separate-agreement requirement for very large model-as-a-service operators and a "Kimi K3" display requirement for very large products. The company also does not release the training data or full recipe, and the Kimi app and API stay closed. So you can run and fine-tune the model, but you cannot fully reproduce how it was made. ### Is Kimi K3 free? The model weights are free to download and run (they went public on July 27, 2026), if you have the hardware, but using K3 through the API is not: it is $3 per million input tokens and $15 per million output, nearly four times the older K2 line's output rate. Third-party providers sometimes offer limited free access, and Moonshot's consumer app has a free tier, but there is no permanent free API tier. ### Can I run Kimi K3 on my own computer? Realistically, no. The weights are public now, but at 2.8 trillion parameters K3 is about 1.5TB to download and needs roughly 1.4TB of GPU memory just to load, so Moonshot recommends a supernode of 64 or more accelerators. That is far beyond any consumer machine: a 512GB Mac Studio does not come close, and one tester who forced it onto a MacBook by streaming experts from disk measured minutes per token. Self-hosting is practical only for organizations with serious GPU infrastructure, usually run as overnight batch jobs rather than interactive chat. For everyone else, the hosted API or a provider like OpenRouter is the route. Our [how to run Kimi K3 locally](https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally) guide covers the routes and hardware in full. ### Is Kimi K3 better than Claude or GPT? On narrow coding and agent benchmarks, K3 competes at or near the top, and it ranks first on a human-preference frontend leaderboard. On general intelligence, independent testing now puts it seventh of 190, behind Claude Opus 5, Fable 5 and GPT-5.6 Sol, and reviewers report it is weaker on non-coding work. It is a strong specialist, not a clear overall winner. Our [Kimi K3 vs Claude](https://geotoolbox.ai/blog/kimi-k3-vs-claude) comparison breaks the head-to-head down by task, cost, and speed. ### Does Kimi K3 hallucinate? Yes, and notably so. On Artificial Analysis's independent testing, K3's hallucination rate rose to 51% from the prior model's 39%, even as its accuracy improved. Higher accuracy came with more confident fabrication, so verify anything that matters before you rely on it. ## Sources - China's Moonshot unveils world's largest open AI model, closing in on US rivals - Reuters, July 2026 - `reuters.com/world/china/chinas-moonshot-unveils-worlds-largest-open-ai-model-closing-us-rivals-2026-07-17` - Kimi K3, and what we can still learn from the pelican benchmark - Simon Willison, July 2026 - `simonwillison.net/2026/Jul/16/kimi-k3` - Kimi K3 - Intelligence, Performance & Price Analysis - Artificial Analysis, July 2026 - `artificialanalysis.ai/models/kimi-k3` - Kimi's open model K3 nears GPT-5.6 Sol and Fable 5 while signaling the end of super cheap Chinese AI - The Decoder, July 2026 - `the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai` - Moonshot AI Releases Kimi K3: A 2.8 Trillion Parameter Open MoE Model - MarkTechPost, July 2026 - `marktechpost.com/2026/07/16/moonshot-ai-releases-kimi-k3-a-2-8-trillion-parameter-open-moe-model-with-kimi-delta-attention-and-1m-context` - Kimi K3 - API Pricing & Benchmarks - OpenRouter, July 2026 - `openrouter.ai/moonshotai/kimi-k3` - Kimi Linear: An Expressive, Efficient Attention Architecture - arXiv 2510.26692 - `arxiv.org/abs/2510.26692` --- ## Best AI Visibility Tools in 2026: 11 Trackers Compared > The 11 best AI visibility tools in 2026, compared on verified 2026 pricing, per-tier engine coverage, accuracy caveats, free options, and cost per prompt. - Canonical: https://geotoolbox.ai/blog/best-ai-visibility-tools - Published: 2026-07-16 · Updated: 2026-08-22 Choosing between [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) tools in 2026 means navigating three problems at once: prices that changed last quarter, vendors grading their own homework, and measurements of an answer engine that rarely says the same thing twice. This comparison takes all three seriously. Engine coverage is reported per tier instead of per marketing claim, and our own bias is disclosed upfront: we build one of these tools. Which tracker you need depends on your team, budget, and how much accuracy skepticism you bring. If you would rather start with the free, manual method before buying, we cover [how to track brand mentions in AI search](https://geotoolbox.ai/blog/track-brand-mentions-in-ai-search) by hand. ## The Best AI Visibility Tools at a Glance Prices and engine coverage below were checked against each vendor's live pricing page in July 2026 (geotoolbox's own row on July 27; Ahrefs Brand Radar re-verified August 2026) - several of these vendors repriced or repackaged within the last quarter.
ToolBest forEngines at entry tierEntry priceFree option
geotoolbox (that's us)Reachability + tracking in one3 (ChatGPT, Perplexity, AI Overviews); all 8 from Pro upFrom $99/mo ($79 annual)7-day trial + free checkers
ProfoundEnterprise depth1 (ChatGPT only); ~10 on Enterprise$99/mo billed yearlyNo trial on Starter
Semrush AI Visibility ToolkitExisting Semrush users5 (ChatGPT, AI Overviews, AI Mode, Gemini, Perplexity)$99/mo per domainFree checker, no trial
Ahrefs Brand RadarBenchmarking on big prompt dataChoose from 5 (Select tier)$199/mo+Free AI checkers
SE VisibleSEO + AI in one stackChatGPT, Gemini, AI Mode, AI Overviews, Perplexity$99/mo10-day trial
Peec AIEuropean teams, agenciesPick 3 of 6; more as add-ons€85/moFree trial
Otterly.aiBudget monitoring4 (AIO, ChatGPT, Perplexity, Copilot)$29/mo ($25 annual)Free trial
RankscaleBudget breadth10 surfaces (incl. Claude, DeepSeek, Mistral)€20/moPro trial
Scrunch AIEnterprise + AI-readable deliveryIn flux (page served 2 lineups when checked)$250/mo7-day trial
AthenaHQAgencies (unlimited seats)9 models$295/mo (~$245 annual)Free credit tier
WritesonicContent teams adding tracking3 platforms$79/mo billed annuallyFree trial
The category goes by several names - AI search visibility tools, LLM visibility tools, AI brand visibility trackers, AI visibility tracking tools - but they all answer one question: when someone asks ChatGPT, Gemini, Perplexity, or Google's AI Overviews about your category, does your brand show up, and who gets cited instead? If you're new to the space, our AI visibility glossary entry covers the fundamentals. With ChatGPT at [900 million weekly users](https://techcrunch.com/2026/02/27/chatgpt-reaches-900m-weekly-active-users/) as of February 2026, that question stopped being optional. ## How We Evaluated (and Why You Should Distrust Every List Like This) Full disclosure: we build geotoolbox, one of the tools on this list. Factor that in when you read our entry. Check the current page-one results for "ai visibility tools" and you'll find a pattern: Profound's list ranks Profound first. Frase's list ranks Frase first. GrowthOS ranks GrowthOS first. Backlinko's list puts Semrush at the top, and Backlinko is owned by Semrush - to its credit it discloses that, though in the footer, after the ranking. The vendor lists carried no conflict-of-interest note at all when we checked in July 2026. The category is young enough that vendor listicles dominate the query, and AI engines cite those same listicles back to you when you ask them for recommendations. So instead of pretending neutrality, we did three things you can check: 1. **Every price was verified against the vendor's live pricing page in July 2026 (Ahrefs Brand Radar re-verified August 2026).** Where a widely repeated number was wrong, we say so - AI answers and half the listicles ranking for this query still quote stale figures. 2. **Engine coverage is reported per tier, not per marketing page.** "Tracks 10 engines" often means "tracks 10 engines on the custom-priced enterprise plan, and exactly one on the plan you can actually buy." 3. **Claims are separated into tested and reported.** Where we could not verify a vendor claim directly, it is attributed, not asserted. Buyers already know this. As Paul Dyer, CEO of the agency /prompt, [put it to Digiday](https://digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/), "If you use three different tools and give them the same prompts, you get three different answers." He's right, and we cover why in the [accuracy section](#how-accurate-are-ai-visibility-tools-really) below. One scope note: this page compares AI visibility trackers specifically. If you want the whole GEO stack (trackers plus content, technical, and authority tools grouped by job), that's our guide to the best generative engine optimization tools. ## Before You Pay: Make Sure AI Can Read Your Site There is a step before tracking that no tool vendor leads with, because there's no subscription in it: check whether AI crawlers can fetch your site at all. GPTBot, ClaudeBot, and PerplexityBot need to fetch your pages to learn what your brand does; on the Google side, Google-Extended is the robots token governing Gemini training and grounding, while AI Overviews and AI Mode ride regular Googlebot. A robots.txt rule copied from a 2023 template, an over-eager WAF or bot-protection layer, or a JavaScript-only render can make you invisible to several engines at once. In our experience running crawler scans, this failure mode is common and expensive: teams pay $300+ a month to track engines whose crawlers their own CDN has been blocking the whole time. The dashboard says "zero mentions" and everyone starts rewriting content, when the actual fix is one firewall rule.
![Three-step AI visibility workflow: check crawler reachability first, then track mentions and share of voice, then act on the gaps.](/blog/best-ai-visibility-tools/fig-reachability-flow.png)
Check reachability before you pay for tracking; act on gaps after you measure them.
Two minutes catches the obvious blockers. Run your domain through our free AI Crawler Checker, which tests real fetches from the major AI crawlers against your live site. If something is blocked, fix that before you spend a dollar on monitoring. We wrote up the full mechanics in our guide to AI crawlers. Reachable? Good. Now the trackers themselves. ## The 11 Best AI Visibility Tools in 2026 Ordered by how we'd shortlist for a typical in-house or agency SEO team, weighing verified cost against engine coverage, whether reachability is covered, and how far the entry tier actually gets you. #1 is ours; the disclosure above applies, and the price-per-prompt table further down puts our number next to everyone else's so you can judge it rather than take our word. Full entries went only to tools with a public, verifiable pricing page - custom-quote-only vendors (Brandlight, Conductor, AIclicks) are out by that rule, not by quality. ### 1. geotoolbox geotoolbox is our AI visibility platform, built around a premise most of the category skips: reachability comes before tracking. Every plan pairs brand-mention and share of voice tracking with a GEO scan, Agent Readiness scan, competitor benchmarking, and AI-traffic analytics (GSC impression views plus GA4 referral behavior), so you see whether engines can read you, whether they mention you, and what the resulting visits do, in one place. Higher tiers add Citation Interceptor, which flags the third-party pages engines cite in your category so you can go earn a mention there, plus content briefs and article generation. See the domain overview for how the pieces fit. **Engines:** 8 total - ChatGPT, Perplexity, Google AI Overviews, Google AI Mode, Gemini, Claude, Bing Copilot, and Grok. Starter runs 3 (ChatGPT, Perplexity, AI Overviews); Scale runs all 8. **Pricing:** per our pricing page, Starter is $99/month ($79/month billed annually) with 50 prompts across three engines and a 7-day free trial. All 8 engines start on the Pro plan at $399/month ($319 billed annually); Scale, at $999/month ($799 billed annually), scales that to 25 brands. **Best for:** teams that want reachability checks, tracking, and AI-traffic analytics on one bill. **Watch for:** Starter tracks a single brand with 50 prompts, and all-8-engine coverage starts on the $399/month Pro plan (150 pooled prompts) - the same entry-versus-full-coverage gap we flag on Profound below. We're also a newer entrant than Semrush or Ahrefs. ### 2. Profound Profound is the enterprise reference point in this category, and the tier gap is the thing to understand before you buy. Per [Profound's pricing page](https://www.tryprofound.com/pricing), the $99/month Starter (billed yearly) tracks ChatGPT only, with 50 prompts. Growth at $399/month adds Perplexity and AI Overviews. The full engine list - around 10, including Gemini, Copilot, Claude, Grok, and DeepSeek - is enterprise-only, custom-priced. What you get for it: browser-level answer collection that mimics real user sessions rather than leaning on APIs, a large conversation-volume dataset, crawler analytics, and agency workspaces. Well-funded and shipping fast. If the budget is enterprise-sized and you need depth plus compliance boxes ticked, it earns the shortlist. **Best for:** enterprises and larger agencies. **Watch for:** the ChatGPT-only entry tier surprising teams who expected multi-engine coverage at $99. ### 3. Semrush AI Visibility Toolkit The strongest "stay in your existing stack" option. The AI Visibility Toolkit is a [$99/month per-domain add-on](https://www.semrush.com/pricing/ai/) (same price monthly or annually) covering ChatGPT, Google AI Overviews, Google AI Mode, Gemini, and Perplexity per Semrush's documentation, with Grok and Claude reserved for enterprise plans. Mentions, quote-level sentiment scoring, share of voice, and citation sources feed directly into the same workspace as your keyword and content data, and a free AI Visibility Checker gives you a one-off snapshot before you commit. If you are choosing the underlying suite first, our [Semrush vs Ahrefs comparison](https://geotoolbox.ai/blog/semrush-vs-ahrefs) tests their core keyword and traffic data on real sites, and our [hands-on Semrush review](https://geotoolbox.ai/blog/semrush-review) covers the full suite's pricing and accuracy. Semrush is also the tool with the most corporate momentum behind it: [Adobe closed its Semrush acquisition](https://news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition) in April 2026 and launched "Adobe Brand Visibility" on Semrush's visibility data that June - which tells you where this category is headed: into the big marketing suites. **Best for:** teams already paying for Semrush. **Watch for:** per-domain pricing stacking up across a portfolio. ### 4. Ahrefs Brand Radar Brand Radar's pitch is scale: 472M+ organic prompts underpinning the benchmarking data. [Two plans](https://ahrefs.com/brand-radar) - Select Platforms from $199/month, where you pick which platforms to track individually (AI Overviews and AI Mode, ChatGPT, Perplexity, Copilot, Gemini), and All Platforms at $699/month, which opens the full 472M+ organic prompt corpus plus 2,500 custom prompt checks a month. YouTube, TikTok, and Reddit tracking are free while in beta on both. That prompt corpus is the differentiator. Instead of only sampling the prompts you configure, you can benchmark against demand-weighted prompt data, which softens the small-sample problem that budget tools have. The trade-off is price and focus: it's a benchmarking instrument more than an action platform, with no crawler-side analysis. **Best for:** brands that want visibility measured against real prompt demand. **Watch for:** widely circulated stale pricing - the $199/month "add-on" figure you'll see in older roundups no longer matches the live pricing page. ### 5. SE Visible SE Ranking spun its AI tracking into a standalone product, [SE Visible](https://visible.seranking.com/), and it's one of the better mid-market values: $99/month entry (Core $189, Plus $355) covering ChatGPT, Gemini, Google AI Mode, AI Overviews, and Perplexity, with a 10-day trial. If you'd rather keep one bill, the classic SE Ranking AI Search add-on route still exists (from around 63 euros/month on top of an SE Ranking plan) and includes the SE Visible dashboard. **Best for:** teams that want SEO and AI visibility side by side without enterprise pricing. **Watch for:** fewer engines and regions than the dedicated enterprise tools; confirm your markets are covered. ### 6. Peec AI Peec is the European favorite: clean UX, daily tracking, and country-specific visibility that most US-built tools handle poorly. [Starter is 85 euros/month](https://peec.ai/pricing) with 50 prompts and one project; Pro is 205 euros (150 prompts), Advanced 425 euros. You pick 3 of 6 engines at entry, with the rest available as paid add-ons - price those before comparing it against flat-rate rivals. Peec is backed by a [$21M Series A](https://techcrunch.com/2025/11/17/as-consumers-ditch-google-for-chatgpt-peec-ai-raises-21m-to-help-brands-adapt/) (November 2025), which lowers the vanish-overnight risk that haunts small tools here, even if funding guarantees nothing about roadmap fit. **Best for:** European teams and agencies tracking multiple countries. **Watch for:** engine add-ons quietly inflating the effective monthly cost. ### 7. Otterly.ai The budget pick that shows up in nearly every credible roundup, for good reason. [Lite is $29/month](https://otterly.ai/pricing/) ($25 billed annually) - the most widely vetted tracker under $30 - covering AI Overviews, ChatGPT, Perplexity, and Copilot at entry, with Gemini, AI Mode, and Claude as add-ons. Standard at $189 and Premium at $489 raise the prompt ceilings. There's a free trial, and the SEO-keyword-to-prompt conversion makes setup quick. **Best for:** SMBs and first-time buyers wanting the safest budget pick. **Watch for:** tight prompt caps at entry, and monitoring-only output - it tells you where you stand, not what to do next. ### 8. Rankscale The widest engine coverage per dollar on this list. [Essentials starts at $20/month](https://rankscale.ai/pricing) and Pro at $99 ($84 billed annually), tracking ten surfaces on every plan - ChatGPT, Perplexity, Google's AI Mode, AI Overviews and Gemini, Claude, DeepSeek, Mistral, Grok, and Copilot - coverage that costs enterprise money elsewhere. You can try Pro free. At this price the trade-offs are predictable: lighter reporting, smaller team features, and less methodological transparency than the big platforms. But if the question is "can I see whether Claude and DeepSeek mention us without a four-figure contract," this is currently the shortest path. **Best for:** budget-conscious teams that specifically need broad engine coverage. **Watch for:** shallow analytics compared to mid-market tools. ### 9. Scrunch AI Scrunch pairs tracking with something no one else on this list does: its Agent Experience Platform (AXP) serves a structured, AI-readable version of your site to crawlers like GPTBot and ClaudeBot without touching your human-facing pages - a real fix for JavaScript-heavy sites that engines parse badly. Pricing here deserves its own caveat, and it proves this article's point about verifying vendor pages: on the day we checked, [Scrunch's pricing page](https://scrunch.com/pricing) was serving two different plan lineups - our browser rendered a $250/month Core plan with four engines and AXP gated to Enterprise, while other same-day checks surfaced a Starter/Growth ladder with broader coverage. What held constant across both: self-serve entry at $250/month, a 7-day trial, SOC 2 compliance, and a custom-priced Enterprise top end. Scrunch has repackaged repeatedly since the acquisition, so confirm the lineup on their page before you commit. [Sitecore acquired Scrunch](https://www.sitecore.com/company/newsroom/press-releases/2026/06/sitecore-acquires-scrunch-to-help-brands-influence-discovery--and-buying-decisions) in June 2026 (Bloomberg reported roughly $225M), positioning it inside a larger DXP play - more consolidation evidence, and worth factoring into a long-term contract decision. **Best for:** enterprise sites with technical rendering problems. **Watch for:** post-acquisition roadmap shifts. ### 10. AthenaHQ A platform agencies like for a structural reason: unlimited seats on every plan. [Starter is $295/month](https://www.athenahq.ai/pricing) (about $245 billed annually) on a credit system, covering 9 models including ChatGPT, Perplexity, AI Overviews, AI Mode, Gemini, Claude, Copilot, and Grok, plus an action center that turns gaps into tasks. A free credit tier and free audit let you test before paying. **Best for:** agencies putting whole teams into one tool. **Watch for:** credit burn - broad prompt sets consume credits faster than the sticker price suggests. ### 11. Writesonic The content-platform route into tracking. Writesonic folds brand monitoring into a writing product at [$79/month billed annually](https://writesonic.com/pricing) ($99 monthly), with deeper layers like sentiment analysis and crawler analytics gated to higher tiers. Self-serve plans track 3 platforms - fine for a content team's first look at AI visibility, thin for serious multi-engine measurement. **Best for:** content teams who want tracking bolted onto production. **Watch for:** mention-frequency trends standing in for precise measurement, and the 3-platform ceiling outside enterprise. ### Also on the Radar Worth knowing but not full entries: **ZipTie** (from $69/month, three engines, strong technical indexation audits), **Hall** (free tier available), and the free one-shot graders - **HubSpot's AEO Grader** (free check across ChatGPT, Perplexity, Gemini; $50/month for continuous tracking), **Semrush's free AI Visibility Checker**, and **Mangools' AI Search Grader** (free, seven models with an account). Free graders are snapshots, not monitoring, but they're the right first step before any subscription. ## LLM SEO Tools: Tracking Is Half the Job Search for **LLM SEO tools** and you'll get lists that mix three genuinely different product types. It's worth separating them, because "we do LLM SEO" can mean any of these: **Trackers** measure LLM visibility directly - brand mentions, citations, sentiment - which is everything above. **Content-side tools** optimize what you publish so models want to cite it; Surfer, Frase, and Clearscope all added AI-citation layers to their content editors in the past year, and our roundup of the best content optimization tools covers that side. **Technical tools** deal with whether models can read you at all: crawler access, rendering, structured markup, and the still-debated llms.txt proposal. Some platforms straddle categories - AthenaHQ and Writesonic bolt content generation onto tracking, Scrunch bolts delivery infrastructure onto it. That's not a flaw, but it explains the confusing pricing spread: you're often paying for a second product category you may not need. Decide which of the three jobs you're hiring LLM SEO tools for first, then compare within that lane. ## How Accurate Are AI Visibility Tools, Really? Here's the section every vendor list skips. The short answer: AI visibility scores are sampled estimates of a random process, and any tool that presents them as rankings is overselling. The numbers are stark. [SparkToro's research](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/) had 600 volunteers run 12 identical prompts 2,961 times across ChatGPT, Claude, and Google's AI: there was less than a 1-in-100 chance of getting the same brand list twice, and roughly 1-in-1,000 of getting the same order. Cross-engine agreement is just as thin - [Kevin Indig's analysis](https://www.growth-memo.com/p/the-consensus-gap) of 3.7 million citations across 20,000 prompts found 91% of cited URLs appear in exactly one engine, and only 2.37% in all three. Strong visibility in one engine transfers poorly to the others - measure each engine separately rather than trusting a blended score. There's also a collection-method gap the sales demos don't mention. Tools gather answers via APIs, via browser sessions that mimic real users, or via hybrid sampling - and logged-in consumer ChatGPT does not answer like the API. Practitioners on Reddit keep rediscovering this: the tool says you're absent, a real signed-in session shows you present, or vice versa. Ask any vendor two questions before buying: how do you collect answers, and how many runs per prompt make up one data point? A third worth adding: whether they capture multi-turn conversations or only the first answer, since real buyers ask follow-ups. What to do with that: 1. **Track frequency, not position.** "Mentioned in 62 of 100 runs" is a real metric; "ranked #3" is noise. SparkToro's own data shows mention frequency stays fairly stable even while individual lists scramble. 2. **Demand sample sizes.** A starter plan sampling 10-25 prompts weekly, once per prompt, is a coin-flip detector. For trend reliability, more runs of fewer prompts usually beats one run of many. 3. **Watch trends across weeks, not day-to-day wobble.** The cost side deserves the same normalization. Entry prices run from $20 to $295 a month, but the prompt allowances behind them spread just as wide (all prices checked July 2026; Ahrefs Brand Radar re-verified August 2026):
ToolEntry price/moPrompts included$ per prompt/moEngines at entry
Rankscale€20varies by plan-10
Otterly.ai$2915 prompts (Lite)~$1.934
geotoolbox$79 (annual)50 prompts/brand~$1.583
Peec AI€8550 prompts~€1.70base set
Profound$99 (yearly)50 prompts~$1.981 (ChatGPT)
Semrush AI Toolkit$99/domain25 tracked prompts~$3.965
AthenaHQ$295credit-based-9
Ahrefs Brand Radar$199+Ahrefs' own prompt corpus (custom prompts on $699 tier: 2,500 checks)-your pick of 5
Read the per-prompt column against the engine column: Profound's $1.98/prompt buys one engine; Peec's 1.70 euros buys three of your choosing. Per-prompt cost only means something multiplied by engines covered and refresh frequency - which is exactly the math vendors' pricing pages make hard. ## How to Choose (by Team, Budget, and Stage) The verdicts, plainly: **Solo operators and SMBs:** Rankscale (€20, priced in euros) if you need engine breadth, Otterly ($29) if you want the most proven budget tracker for brand mentions, geotoolbox Starter ($79 annual, after a 7-day free trial) if you want reachability checks and AI-traffic analytics in the same bill. All three cost less per month than one hour of consulting. **Agencies:** the deciding features are seats and client workspaces, not engines. AthenaHQ's unlimited seats, Peec's project structure, and Profound's agency mode are the three built for this. Add per-client costs carefully - per-domain pricing (Semrush) multiplies fast across a book of clients. **Enterprise:** Profound and Scrunch for depth and compliance, Ahrefs Brand Radar for demand-weighted benchmarking, Semrush-Adobe if procurement already loves the suite. This is also where doing it in-house stops being crazy - see our breakdown of GEO services vs software for when to buy tooling versus hire the job out, and the best AI SEO agencies guide if you decide to hire. Whatever the segment, run the same three TCO checks before signing: (1) which engines are add-ons rather than included, (2) where the prompt or credit wall sits and what the next tier costs, and (3) how often data refreshes - weekly refresh at a daily-refresh price is a quiet 7x difference in data volume. ## What AI Visibility Tools Cannot Do No tracker, at any price, can guarantee you a mention, edit what a model already learned, or make sampled measurements behave like deterministic rankings. The same prompt can produce different answers for different users in the same hour. Anyone selling certainty in this category is selling past what the technology does. What tracking buys you is a map of absence: which prompts you're missing from, who gets cited instead, and which sources feed those citations. Acting on the map is a separate job - publishing citable content, earning brand mentions on the pages engines pull from, and keeping your site technically readable to crawlers. We walk through that loop in how to track AI visibility, and how to turn the raw counts into a comparable metric in our AI visibility score guide. Buy the dashboard for the map. Budget separately for the fixing. ## Frequently Asked Questions ### What is the best AI visibility tool? There is no single best - and any list that says otherwise is usually ranking its own product (including, arguably, this one; we build geotoolbox). The short version: match the tool to your segment and budget, then run the three TCO checks from the [how-to-choose section](#how-to-choose-by-team-budget-and-stage) before signing anything. ### How much do AI visibility tools cost in 2026? Free one-shot graders at $0; real monitoring starts low (Rankscale at €20, Otterly at $29, geotoolbox at $79 annual); mid-market runs roughly $85-295 (Peec's 85-euro entry, Semrush's add-on, Scrunch, AthenaHQ); Ahrefs Brand Radar spans the range and beyond ($199/mo per platform on the Select tier, up to $699/mo for all platforms), and custom Profound tiers price by quote. Verified against vendor pricing pages, July 2026 (Ahrefs Brand Radar re-verified August 2026). ### Are there free AI visibility tools? Yes, three kinds: one-shot graders (HubSpot AEO Grader, Semrush AI Visibility Checker, Mangools AI Search Grader), free tiers or trials of paid tools (geotoolbox's 7-day trial, Hall, AthenaHQ credits), and DIY - a GA4 referral filter catches traffic clicking through from ChatGPT and Perplexity, and costs nothing. ### How many prompts do you need to track? Fewer than vendors imply, sampled more often. AI answers are random enough that repeated runs of a focused prompt set usually beat one run of a big list - balance breadth of coverage against repeatability. Prioritize run frequency and mention-rate trends over prompt-list size, and treat week-over-week movement as signal only when it persists. ### Does Google Search Console show AI traffic now? Partially. Google launched [generative AI performance reports](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports) in Search Console on June 3, 2026 - dedicated impression views for AI Overviews and AI Mode. It's rolling out to a subset of sites first, covers impressions rather than clicks, and only Google surfaces. For ChatGPT, Perplexity, and the rest you still need the tools above. ## Where to Start Not with a subscription. Run the free checks first: confirm AI crawlers can reach your site, then grab a one-shot visibility snapshot to see where you stand. Ten minutes, zero dollars, and you'll know whether your problem is reachability, visibility, or neither. If the snapshot shows gaps worth tracking, pick the tier that matches how you'll act on the data - and if you want the reachability check, the tracking, and the AI-traffic analytics in one place, that's the exact gap we built geotoolbox to fill. Start with the free AI readiness check and see what the engines see. ## Sources - SparkToro: "AIs are highly inconsistent when recommending brands or products" - Rand Fishkin, January 2026 - `sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility` - Growth Memo: "The Consensus Gap" - Kevin Indig, May 2026 - `growth-memo.com/p/the-consensus-gap` - TechCrunch: "ChatGPT reaches 900M weekly active users" - February 2026 - `techcrunch.com/2026/02/27/chatgpt-reaches-900m-weekly-active-users` - Google Search Central: "Introducing Search Generative AI performance reports in Search Console" - June 2026 - `developers.google.com/search/blog/2026/06/gen-ai-performance-reports` - Digiday: "Marketers question expensive AI visibility tools as inconsistent results fuel skepticism" - May 2026 - `digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism` - Adobe newsroom: "Adobe completes Semrush acquisition" - April 2026 - `news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition` - TechCrunch: "Peec AI raises $21M Series A" - November 2025 - `techcrunch.com/2025/11/17/as-consumers-ditch-google-for-chatgpt-peec-ai-raises-21m-to-help-brands-adapt` - Sitecore newsroom: "Sitecore acquires Scrunch" - June 2026 - `sitecore.com/company/newsroom/press-releases/2026/06/sitecore-acquires-scrunch-to-help-brands-influence-discovery--and-buying-decisions` Every source above is also hyperlinked inline where its claim appears in the article. Tool pricing and engine coverage were verified against each vendor's live pricing page in July 2026 (Ahrefs Brand Radar re-verified August 2026). --- ## What Is Claude Code? Anthropic's Agentic Coding Tool, Explained > What is Claude Code? Anthropic's agentic coding tool: it reads your codebase, edits files, and runs commands from plain English. What it costs and how it works in 2026. - Canonical: https://geotoolbox.ai/blog/what-is-claude-code - Published: 2026-07-16 · Updated: 2026-07-22 Claude Code is Anthropic's agentic coding tool, and it is not the Claude you chat with. That distinction trips up almost everyone who searches for it, including Google, whose top organic result for "claude code" is currently the sign-in page for the chat assistant. It also moves faster than the answers about it. The install command Google's own AI summary still hands you is deprecated. The "terminal-only" framing is a version behind. Straighten both out before you decide whether the tool is for you. ## What Claude Code Is Claude Code is Anthropic's agentic coding tool. It reads your codebase, edits files, runs commands and tests, and carries out multi-step development tasks from plain-English instructions, under permissions you control. The verbs are the part people miss. Claude Code does not suggest code for you to copy. It does the work: opens files, makes changes across them, runs the test suite, reads what broke, and tries again. **Two things it is not**, and both confusions are common. **It is not a model.** People say "I used Claude Code" the way they say "I used Opus," and ask which is better. The question does not parse. Claude Code is the harness; Opus and Sonnet are the models that run inside it. You pick the model with a `/model` command, and the same agentic loop runs underneath whichever one you choose. **It is not chat with access to your files.** A chat assistant that can read your repo still hands the work back to you. Claude Code executes. It has a shell, it can commit to git, and it can run a build and read the error output. That is the difference between a tool that describes a fix and one that applies it, and it is what the word [agentic AI](https://geotoolbox.ai/blog/agentic-ai) is doing in the definition. The practical shape of it: you open a terminal in a project, type `claude`, and describe what you want in a sentence. It explores the codebase itself rather than waiting for you to paste the relevant files in. For a working definition of the broader category, our [AI agent](https://geotoolbox.ai/glossary/ai-agent) glossary entry covers the general pattern; Claude Code is one of the more literal implementations of it. ## Claude Code vs Claude: Not the Same Thing Claude is the assistant you talk to. Claude Code is the tool that does the work on your machine. Same models underneath, different products. Google is not helping. When we pulled the US results for "claude code" on 15 July 2026, the number one organic result was the sign-in page for [Claude AI](https://geotoolbox.ai/blog/what-is-claude-ai), the chat assistant. Not the coding tool. The top related query on Google Trends was "claude vs claude code," and Google's own People Also Ask box drifted to "Is Claude better than ChatGPT?" That is a search engine trying to answer a question about a developer tool with information about a chatbot. The split is clean once you see it. **Claude** answers you, in a browser, a desktop app, or on your phone. **Claude Code** runs where your code lives, and it changes things. What differs is the harness, not the intelligence inside it, which is why [how Claude works](https://geotoolbox.ai/blog/how-does-claude-work) under the hood is the same story in both products, and why the answer to "which is better" is neither. They are not competing. ## How Claude Code Works: the Agentic Loop Everything Claude Code does reduces to one loop: look at the project, decide on an approach, make a change, run something, read the result, correct. Three parts make it work, and one mode decides when it starts.
![The Claude Code agentic loop: gather context, decide, edit, run, read the output, correct, repeat.](/blog/what-is-claude-code/agentic-loop.png)
The return path is what makes it a loop: it reads command output and test results to decide whether to go again.
### It Gathers Its Own Context You do not assemble the context. Claude Code searches the project, opens the files it decides are relevant, and reads them. That is why it behaves differently from a chat window, where the quality of the answer is capped by how well you chose what to paste. The cost is that everything it reads occupies the same budget. Your instructions, the files it opened, every command's output, and the entire conversation share one [context window](https://geotoolbox.ai/glossary/context-window), and it fills faster than most people expect. We covered the mechanics of that in detail in our guide to the [Claude Code context window](https://geotoolbox.ai/blog/claude-code-context-window), including how to read the meter before output quality starts degrading. ### It Runs Things and Reads What Happens The loop closes because Claude Code can execute. It runs the test, reads the failure, edits the file, runs the test again. A chat assistant guesses whether its fix worked. Claude Code checks. ### It Uses Your Tools Shell commands, git, your test runner, and any MCP servers you have connected are available to it, within the permissions you grant. Connecting a tool is what turns a general coding assistant into something that can open a pull request or query your database, and the [open-source tooling around Claude Code](https://geotoolbox.ai/blog/claude-code-open-source-tools) has grown up around exactly that seam. ### Plan Mode Most guides skip this, and it is the feature that changes how the tool feels. In plan mode, Claude Code investigates and proposes an approach without touching anything. You read the plan, correct the parts it got wrong, and only then let it execute. Use it by default on anything non-trivial. The failure mode of an agentic tool is not usually a bad line of code. It is confidently doing the wrong task well, and a plan you can reject costs a few seconds to prevent exactly that. ## Where You Can Actually Use It Claude Code is not terminal-only. That framing was true once, it still shows up in guides that rank for this topic, and it is out of date. Anthropic's [own overview](https://code.claude.com/docs/en/overview) puts it plainly: Claude Code "runs on several surfaces: the terminal, IDE extensions, a desktop app, and the web." The [how-it-works page](https://code.claude.com/docs/en/how-claude-code-works) goes further, listing access through the terminal, the desktop app, IDE extensions, claude.ai/code, Remote Control, Slack, and CI/CD pipelines. Notably, the docs never commit to a number, and different pages enumerate different lists, so a guide that hands you a firm count is inventing precision the docs never claim. Three details are easy to get wrong, and all three are current as of July 2026: The **desktop app** is a real Claude Code surface, which several AI assistants will tell you it is not. Anthropic's product FAQ answers the question directly: "Yes. Max, Pro, Team, and Enterprise users can access Claude Code on the Claude desktop app." The **JetBrains** plugin is still labeled [Claude Code [Beta]](https://plugins.jetbrains.com/plugin/27310-claude-code-beta-) on the JetBrains Marketplace, by Anthropic itself. The VS Code extension carries no such label. If you work in an IDE, that distinction is worth knowing rather than discovering. **Remote Control** is the one people do not expect: your code and the execution stay on your machine while you drive the session from a browser somewhere else. The important part is that the loop does not change. Anthropic is explicit that the interface determines how you see and interact with Claude, while the underlying agentic loop stays identical. Choosing a surface is a preference, not a capability decision. ## How to Install Claude Code Use the native installer. Anthropic labels it "Native Install (Recommended)" in the [setup docs](https://code.claude.com/docs/en/setup). Pick the line for your platform. macOS, Linux, or WSL: ```bash curl -fsSL https://claude.ai/install.sh | bash ``` Windows, in PowerShell: ```powershell irm https://claude.ai/install.ps1 | iex ``` Or through a package manager, if you already use one. On macOS with Homebrew: ```bash brew install --cask claude-code ``` On Windows with winget, in PowerShell: ```powershell winget install Anthropic.ClaudeCode ``` Then point it at a project and start: ```bash cd your-project claude ``` You may have read that you install Claude Code with `npm install -g @anthropic-ai/claude-code`. That command still exists and the setup docs still document it, so it is not wrong exactly. It is just no longer the path Anthropic points you at. The [project's own README](https://github.com/anthropics/claude-code) is direct about it: "Installation via npm is deprecated. Use one of the recommended methods below." The practical reason to care is Node. The npm package requires Node.js 22 or later as of v2.1.198, though the docs note that on an older version npm only prints a warning and the install still works. The native installer skips the question entirely. It wants 4 GB of RAM and a reasonably current OS, and Anthropic's [setup docs](https://code.claude.com/docs/en/setup) list the supported versions in full. If you have been putting off installing Claude Code because you did not want to manage a Node toolchain for it, that reason expired. Anthropic marks one path recommended and the other deprecated, and Google's AI summary still hands you the deprecated one. That gap is not an install problem. ## What Claude Code Costs No, Claude Code is not free. It needs either a paid Claude subscription or an Anthropic Console account, and neither has a free tier. There are two ways to pay, and mixing them up is a genuine source of confusion. A subscription buys you a bounded allowance across Claude. A Console account bills you per token, and Anthropic charges no separate Claude Code fee on top: it "consumes API tokens at standard API pricing."
PlanPriceClaude Code access
Pro$20/month, or $17/month billed annually ($200 up front)Included
Max 5x$100/monthIncluded, for larger codebases
Max 20x$200/monthIncluded, most access to Claude models
Console (API)Standard API pricing, per tokenNo separate Claude Code fee
Fast mode$10 / $50 per million tokens (in, out)Research preview, consumption plans
Figures verified against Anthropic's [Claude Code product page](https://claude.com/product/claude-code) on 16 July 2026. Pricing on this product has moved more than once in the past year, so check before you budget on it. Our [Claude pricing](https://geotoolbox.ai/blog/claude-pricing) guide tracks the full plan lineup, including Team and Enterprise seats. The sticker price is the easy part. What almost nobody publishes is what a working day costs, and Anthropic does: across enterprise deployments the average is **around $13 per developer per active day**, and **$150 to $250 per developer per month**, with costs staying under $30 per active day for 90% of users. Those numbers are the useful anchor. If you are on Max 20x at $200 a month and you code most days, you are roughly at the enterprise average and the flat rate is doing its job. If you only code a couple of days a week, that same per-day figure lands nearer $110 a month, which is below Max 20x and close to Max 5x, and the case for paying per token on a Console account rather than buying a subscription tier you will not fill. Treat the number as a rough ceiling rather than a quote: it is an enterprise deployment figure, and individual usage diverges hard in both directions. And if your bill is running well past $30 on an active day, that is not the price of the tool, it is a usage pattern you can change. The swing is mostly context, not the plan you picked, and our guide to [reducing Claude Code token costs](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) covers the habits that move it. ## The Usage Limits Everyone Argues About This is the complaint. Not the price, the limits. r/ClaudeAI runs a [standing megathread for Claude usage limits](https://www.reddit.com/r/ClaudeAI/comments/1s7fcjf/claude_usage_limits_discussion_megathread_ongoing/) that has collected over 5,800 comments, and a subreddit does not maintain a permanent thread for a topic that comes up occasionally. Read the thread and the grievance sharpens: consumption is unpredictable. You cannot see what a request will cost before you make it, you cannot easily tell which habit burns the allowance, and the limit tends to arrive as a surprise mid-task rather than as something you saw coming. The striking part is that Anthropic does not really dispute this. Its own [cost documentation](https://code.claude.com/docs/en/costs) hedges every number it gives you: the dollar figure in `/usage` "is an estimate," the figures "are approximate" and exclude usage from other devices, and costs "vary widely" between developers. When the vendor's own docs decline to promise you an accurate number, the community is not imagining the problem. What is documented is the shape. Usage runs against a five-hour session limit with a weekly limit on top, both documented for [Pro and Max](https://support.claude.com/en/articles/9797557-usage-limit-best-practices). Crucially, subscription capacity is shared with the rest of Claude rather than being a coding-only pool, so a heavy afternoon in the chat window is spending the same allowance your evening coding session needs. What is not documented is any specific number. You will find confident figures elsewhere, and several AI assistants will quote you exact prompt counts per plan. We could not verify any of them against a primary source, so we are not repeating them. The mechanism underneath is real, though. A July 2026 logging-proxy test by the consultancy Systima, [posted to Hacker News](https://news.ycombinator.com/item?id=48883275), measured Claude Code sending around 33,000 tokens before it even reads your prompt, against roughly 7,000 for OpenCode. Treat it as one un-peer-reviewed test by a firm with a product in the space rather than a settled fact, but the direction is consistent with the docs: the loop reads a lot before it acts, and reading is what you are paying for. ## Is Claude Code Worth It in 2026? For most working developers, yes, with the caveat that it rewards a particular way of working. The adoption is not marketing. [A 2026 AI tooling survey](https://newsletter.pragmaticengineer.com/p/ai-tooling-2026) published on Gergely Orosz's The Pragmatic Engineer in March 2026, with research by Elin Nilsson, found Claude Code went from nothing to the most-used AI coding tool among the developers they surveyed in about eight months, and the most-loved one. That is a real signal from 906 working engineers, most of whom use several of these tools side by side. Now the part the enthusiastic guides leave out. Search interest in Claude Code peaked around February and March 2026 and has come down since. Our own pull of Google search volume on 15 July 2026 puts it at roughly 55% of its March peak, down about a third quarter on quarter. That looks like a tool cooling off, and it is not. The other agentic AI tools we tracked fell harder: OpenClaw is down to 25% of its search peak and Antigravity to 37%, against Claude Code's 55%. Claude Code held more of its peak than any competing tool we measured, and is still up around 124% year on year. This is a category coming down off a hype spike together, not one product losing to another. The read we would give a colleague: the tool works, the market has stopped being amazed that it works, and what is left is the ordinary question of whether it fits your workflow. If you work in a codebase most days and you are willing to use plan mode and watch your context, it earns its subscription quickly. If you want a chat window that writes functions on request, you will pay for a loop you are not using. If you do install it, start with plan mode on anything larger than a one-file change, and keep an eye on the context window, because that is what governs both output quality and cost. ## What the AI Engines Say about Claude Code Plenty of people never read an article like this one. They ask an assistant "what is Claude Code" and take the answer. So what is that answer made of? We measured it. "Claude code" carries roughly 49,300 monthly AI searches in the US, and when we pulled the sources that ChatGPT draws on for that topic in July 2026, the ranking was not what you would expect:
Source domainTimes cited
reddit.com220
anthropic.com155
github.com102
cursor.com64
businessinsider.com59
Reddit outranks Anthropic's own site as a source about Anthropic's own product. The vendor is not the primary voice describing the vendor. Which brings back the install command. Google's AI Overview for "what is claude code" currently walks readers through getting started with the npm command, the one Anthropic's repo marks deprecated. Its definition says the tool runs "directly inside your terminal or IDE," which the docs outgrew. Here is the part that should worry you. Look at what that Overview actually cites. Anthropic's own overview page, cited three times, says "Native Install (Recommended)" and never mentions npm. Its most-cited independent source, builder.io, cited three times as well, tells readers outright that the native installer replaced npm. The npm command the Overview leads with traces to a different citation: a Zapier tutorial from August 2025, cited twice, that still presents `npm install -g @anthropic-ai/claude-code` as the way to install Claude Code, with no note that it has since been deprecated. And the one Anthropic page that does still document npm, the setup docs, is not in the citation set at all. So the engine did not lack for better sources. It cited them, and it led with the stale one anyway. We watched a sharper version of the same thing. We loaded twelve sources into a research model, three of them Anthropic's own pages: the Claude Code overview, the how-it-works page, and the product page. We asked it to check the install method. It reported that the docs teach `npm install` and attributed that to the official documentation. None of those three official pages contains the string `npm install` at all. The command was real, but it came from the third-party tutorials in the set, two of which flag it as outdated in the same breath. The model lifted it from those, pinned it on Anthropic's docs, and dropped the deprecation caveat on the way. It was not summarizing a stale source. It was blending the set and crediting the wrong one. To be clear about credit: independent guides had already spotted the npm drift before we did, and several are among the sources the engines cite. Publishing the correction was not enough, because the engine had the corrected sources in hand and still surfaced the stale one. Engines lag hardest on facts that changed recently, which is exactly when you cannot afford to wait for them to catch up. The weighting is not yours to fix by writing a better page. That is the whole problem geotoolbox exists to measure, and Claude Code is just the worked example. If an engine can do this to Anthropic, on a page Anthropic controls, it can do it to you, and your analytics will not show it. Our guide on how to [get cited in Claude](https://geotoolbox.ai/blog/claude-seo) covers the mechanics of ending up in that citation set rather than outside it. ## Being Right in Your Docs Is Not Enough Anthropic publishes accurate, current documentation about its own product and labels the recommended path plainly, and the answer Google hands people who ask what Claude Code is still leads with the deprecated npm command. The gap was never in the documentation. It was in what the engine did with it. Your product has the same exposure, without Anthropic's docs traffic to soften it. If an engine has been describing your pricing, your features, or your setup from a guide someone wrote eighteen months ago, that is invisible from where you are standing. That gap is what geotoolbox was built to close. Our [citation interceptor](https://geotoolbox.ai/features/citation-interceptor) shows you which sources the engines actually reach for when they describe you, so you can go and fix the wrong ones rather than guessing which they are. ## Frequently Asked Questions ### Is Claude Code free to use? No. It requires either a paid Claude subscription, starting with Pro at $20 a month, or an Anthropic Console account billed per token. There is no free tier. On the Console path Anthropic charges no separate Claude Code fee, so you pay standard API rates for what you use. ### Is Claude Code better than ChatGPT? They are different kinds of thing, so the comparison usually means something else. ChatGPT is an assistant you talk to; Claude Code is an agentic tool that edits and runs your code. The fair comparison is Claude Code against other agentic coding tools like Cursor or OpenAI's Codex, not against a chat product. ### Does Claude Code work offline? No. The models run on Anthropic's servers, not on your machine, so every request needs a network connection. What runs locally is the harness: the file reads, the shell commands, and the edits all happen on your machine, which is why it can work on a private codebase without uploading the whole repository. ### Does Anthropic train on my code? It depends on how you pay. Under commercial terms, which covers the Console API, Team, and Enterprise plans, Anthropic's [data usage docs](https://code.claude.com/docs/en/data-usage) state it does not train its models on the code or prompts you send through Claude Code unless you explicitly opt in. On the consumer plans, Free, Pro, and Max, your data can be used to improve future models when that setting is turned on, and you can turn it off at any time in your privacy settings. If keeping your code out of training is a hard requirement, a Console, Team, or Enterprise plan is the safer default. ### Do I need to know how to code to use Claude Code? It helps considerably, but the tool is not limited to writing software. A substantial part of the community uses it as a general local automation agent for file operations, data wrangling, and research. The limit is review: if you cannot evaluate what it did, you cannot catch it confidently doing the wrong thing. ### Is Claude Code the same as the Claude Agent SDK? No, though they are closely related. The Agent SDK is Claude Code packaged as a library, so you can build your own agent on the same harness rather than using Anthropic's interface. Claude Code is the product; the Agent SDK is the engine offered separately. ## Sources - Claude Code overview (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/overview` - Set up Claude Code (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/setup` - How Claude Code works (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/how-claude-code-works` - Manage costs effectively (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/costs` - Data usage (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/data-usage` - Usage limit best practices (Anthropic, Claude Support) - `support.claude.com/en/articles/9797557-usage-limit-best-practices` - Claude Code product page (Anthropic) - `claude.com/product/claude-code` - anthropics/claude-code (Anthropic, GitHub) - `github.com/anthropics/claude-code` - Claude Code [Beta] plugin listing (Anthropic, via JetBrains Marketplace) - `plugins.jetbrains.com/plugin/27310-claude-code-beta-` - AI tooling in 2026 (Gergely Orosz and Elin Nilsson, The Pragmatic Engineer, March 2026) - `newsletter.pragmaticengineer.com/p/ai-tooling-2026` - Claude Usage Limits Discussion Megathread (r/ClaudeAI) - `reddit.com/r/ClaudeAI/comments/1s7fcjf/claude_usage_limits_discussion_megathread_ongoing` - Claude Code sends 33k tokens before reading the prompt; OpenCode sends 7k (Hacker News, July 2026) - `news.ycombinator.com/item?id=48883275` --- ## Best Perplexity Rank Trackers (2026): An Honest Comparison > The best Perplexity rank trackers in 2026, compared honestly: what each tracks, which ones gate Perplexity behind a higher tier, the free way to check, and how to pick. - Canonical: https://geotoolbox.ai/blog/best-perplexity-rank-tracker - Published: 2026-07-13 · Updated: 2026-08-22 Perplexity is worth the attention: one [Seer Interactive analysis](https://www.seerinteractive.com/insights/case-study-6-learnings-about-how-traffic-from-chatgpt-converts) of a single client found its referral traffic converted at about 10.5 percent, against roughly 1.76 percent for Google organic (ChatGPT ran higher still, at 15.9 percent). So the demand for a good Perplexity rank tracker is real, and so is the noise around it. Search "best Perplexity rank tracker" and you get a wall of lists that each crown the author's own tool. This one is grouped by what the tools actually do, tells you where the free options end, and says plainly which one is ours. Geotoolbox, the site you are reading this on, makes an AI-visibility tool. We have included it below, placed it in the category it actually fits, and not ranked it first. Pricing and engine coverage were checked against each vendor's own pages in July 2026 (AthenaHQ, Otterly, Peec AI, and Ahrefs Brand Radar re-verified August 2026), and the Perplexity mechanics against Perplexity's crawler docs and peer-reviewed measurement research. Pricing changes often, so confirm the current number on each vendor's site before you buy. ## What a Perplexity Rank Tracker Actually Tracks Start with the thing most tool pages skip: Perplexity does not rank pages one through ten. It runs a live search, reads a handful of pages, and synthesizes an answer that cites roughly three to five of them. A [study of 21,143 citations across ChatGPT, Google, and Perplexity](https://arxiv.org/abs/2604.25707) found Perplexity cites more sources per answer than ChatGPT, but there is still no fixed "position 1." So a "Perplexity rank tracker" is not tracking a rank. It is tracking whether you show up in the cited set, and how often. That splits into two things a tool measures, and you should know which you are buying: - **Citations** are the numbered, clickable source links under an answer. They drive referral traffic. - **Mentions** are when Perplexity names your brand in the text with no link. They build awareness but send no clicks. A good tracker separates the two, because "we got mentioned" and "we got the click" are different outcomes. On top of that, the useful ones report your **citation rate** (how often you appear for a prompt set), [**share of voice**](https://geotoolbox.ai/glossary/share-of-voice) against named competitors, answer **sentiment**, which specific **URLs** earn the citation, who **replaces you** when you are absent, and how much the results **move** run to run. If a dashboard shows a single visibility score with no method behind it, treat that as a red flag, not a metric. What no tracker can give you is a guaranteed placement or a stable leaderboard. Perplexity's cited set shifts every run. That is how the product works, and no tracker changes it. For the broader picture of measuring this across engines, we cover the [full AI-visibility measurement stack](https://geotoolbox.ai/blog/how-to-track-ai-visibility) separately; here the focus is Perplexity and the tools built for it. ## How Perplexity Rank Tracking Differs from Google Rank Tracking Three differences change what you should expect from a tool. **It searches live, and it favors fresh pages.** Unlike a model answering from training data, [Perplexity](https://geotoolbox.ai/blog/what-is-perplexity) fetches current pages for most answers and leans hard on recency. A page updated last month can displace one that has ranked on Google for years. The freshest clear answer wins, whatever its Google position. **The same prompt gives different answers.** Ask twice and the cited sources can change by time, location, personalization, and which model variant ran. That is how Perplexity works, and a [statistical study of AI-search visibility](https://arxiv.org/abs/2603.08924) put numbers on it: citation results behave like samples from a distribution across repeated runs, so a single check gives a misleadingly precise picture. A tracker earns its price by sampling a prompt set on a schedule and reporting the trend, rather than telling you where you sit at 3pm on a Tuesday. It is also why the answer to a query [fans out into many sub-answers](https://geotoolbox.ai/blog/query-fan-out) that no single lookup captures. **Ranking on Google does not carry over.** Plenty of pages that own the Google result never get cited by Perplexity, because Perplexity has to be able to lift a clean, attributable claim from the page. A strong ranking with the answer buried under three paragraphs of throat-clearing loses to a shorter page that states the fact in the first line. Which leads to the check almost every tool skips. ## Check This Before You Buy Any Tracker: Can PerplexityBot Even Reach You? Here is the trap. A tracker reports "you are not cited." You assume it is a content problem and start rewriting. But there is a step before content: Perplexity has to be able to reach the page at all. Perplexity uses two agents, and its [official crawler documentation](https://docs.perplexity.ai/guides/bots) names both. **PerplexityBot** is the crawler that surfaces and links pages in Perplexity's search results. **Perplexity-User** is the agent that visits a page in real time when someone asks a question. If your robots.txt blocks PerplexityBot, if a bot-protection rule or a WAF challenges it, if the answer only renders after JavaScript, or if the page is slow to respond, you can be doing everything right on content and still be invisible. There is also a gap between being fetched and being cited: a page can be crawled and still never make the answer. In our experience, this is where a surprising number of "we are not showing up in AI" cases actually resolve, and it is the cheapest thing to rule out first, because a paid tracker cannot fix a reachability block, it can only keep reporting the symptom. Confirm the bots can fetch and render your key pages before you spend on tracking. You can do that in a minute with our free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker), which tests whether [PerplexityBot](https://geotoolbox.ai/glossary/perplexitybot) and the other AI agents can actually load your page. One caveat from Perplexity's own docs: Perplexity-User "generally ignores robots.txt," so a robots block stops the indexing crawler (PerplexityBot) but not the live fetcher, and it is not a privacy control.
![Track Perplexity in the right order: step 1 check reachability so PerplexityBot and Perplexity-User can fetch your pages, step 2 track citations across engines on a schedule, step 3 act on what moves.](/blog/best-perplexity-rank-tracker/reachability-first-perplexity-tracking.png)
Reachability first: a tracker can only report the symptom until the crawler can reach the page.
## The Best Perplexity Rank Trackers, Compared One thing to know before the list: many of the guides you will find are published by a tracker vendor, and the vendor-run ones reliably rank their own product at number one. So here is our disclosure up front, and geotoolbox is not sitting at the top of this table. The tools are grouped by what they actually do, with the one detail buyers get burned on called out: whether Perplexity is in the plan you are pricing, or gated behind a much higher tier.
ToolCategoryPerplexity includedFree optionPaid fromBest for
Otterly.aiDedicated trackerBase planTrial~$29/moBudget multi-engine monitoring
Keyword.comDedicated trackerBase planTrial~$24.50/moLowest-cost, model-specific checks
Peec AIDedicated trackerBase planTrial~€85/moCompetitive benchmarking
RankabilityDedicated trackerBase planTrial~$99/moSource-level citation detail
ZipTieDedicated trackerBase planTrial~$69/moTurning tracking into action
NightwatchDedicated trackerBase plan14-day trial~€79/moPairing Google ranks with AI citations
Scrunch AIEnterprise trackerBase plan7-day trial~$250/moMulti-brand / enterprise
AthenaHQEnterprise trackerBase planFree audit~$295/mo ($245 annual)Self-serve, no sales call
ProfoundEnterprise tracker$399 GrowthFree report$99 (ChatGPT); $399 for PerplexityLarge-team citation analytics
Semrush AI VisibilitySEO-suite add-onBase planFree checker~$99/moExisting Semrush users
SE RankingSEO-suite add-onAdd-onTrial~$89/mo add-onAgencies wanting white-label
Ahrefs Brand RadarSEO-suite add-onAdd-onFree checker$199+/platform ($699 all)Existing Ahrefs users
HubSpot AEO GraderFree graderOne-off gradeFree~$50/mo ongoingA one-off snapshot
GeotoolboxReachability + trackingStarter tier7-day trial + free toolsFrom $99/mo ($79 annual)Fixing why you are not cited, then tracking
*Pricing checked July 2026; AthenaHQ, Otterly, Peec AI, and Ahrefs Brand Radar re-verified August 2026. Confirm the current tier on each vendor's site.* ### Dedicated Perplexity Trackers These run prompt sets across AI engines and report where you get cited. Otterly.ai is a popular budget pick, with plans from around $29 a month and Perplexity in the base tier; the lever on price is how many prompts you track. Keyword.com is the lowest-cost dedicated option and can check specific Perplexity model variants, which matters if you care about Sonar versus Pro results. Peec AI leans into competitor benchmarking and multilingual coverage from about €85 a month. Rankability is strong on source-level detail, showing the exact URLs Perplexity cites, at $99 a month with Perplexity in the base plan. ZipTie ($69) puts more weight on telling you what to change, not just what happened. Nightwatch pairs long-standing Google rank tracking with AI visibility and keeps Perplexity on its entry tier from about €79 a month, so it suits teams that want AI citations sitting next to their classic rankings. ### Enterprise Trackers Scrunch AI targets multi-brand and enterprise monitoring from around $250 a month. AthenaHQ is self-serve from roughly $295 a month, or about $245 billed annually, with a free audit so you can see your standing before paying. Profound is the name buyers ask about most, and it is where the tier structure matters: its entry plan starts at $99 but tracks ChatGPT only, and Perplexity arrives on the $399 Growth tier (published on its pricing page as of July 2026). If Perplexity is your priority, price the plan that actually includes it, not the headline number. ### SEO Suites With an AI Add-On If you already pay for a suite, you may not need a separate subscription. Semrush AI Visibility tracks Perplexity alongside the other engines from about $99 a month standalone and plugs into the keyword data you already use. SE Ranking's AI Search add-on runs about $89 a month on top of a base plan and brings the white-label reporting agencies like; its standalone SE Visible product runs closer to $189. Ahrefs Brand Radar ties AI mentions to Ahrefs' backlink data, but the AI layer runs about $199 a month per platform for selected platforms or $699 for all of them, on top of a base Ahrefs plan, so the all-in climbs quickly and it makes sense mainly if you already live in Ahrefs. For the wider set of tools that span more than Perplexity, see our [best GEO tools guide](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools). ### Where Geotoolbox Fits (and Where It Doesn't) Ours is not a Perplexity specialist. Geotoolbox tracks Perplexity as one of eight engines a single scan covers, so it fits teams who want Perplexity in context with ChatGPT, Gemini, Google AI Overviews and the rest rather than in isolation. And unlike Profound and the Ahrefs and SE Ranking add-ons, Perplexity is included at the entry price: it is one of the three engines on the [$99 Starter tier](https://geotoolbox.ai/pricing) ($79 billed annually, after a 7-day free trial), where Profound, for comparison, gates Perplexity behind its $399 Growth plan. The full eight-engine set sits on higher tiers. Where it differs from most trackers is that it starts one step earlier. The same account checks whether the AI crawlers can reach your pages, then a [GEO Scan](https://geotoolbox.ai/features/geo-scan) shows whether each engine cites you and gives a 0 to 100 visibility score, with a running [brand-monitoring dashboard](https://geotoolbox.ai/features/domain-overview) once you track over time. If your problem is "why am I not cited," that reachability-first order is the point. ## Is There a Free Way to Track Perplexity? Yes, with a caveat: the free methods are manual, and most "free" tools are one-off snapshots, not tracking. The genuinely free method is your own analytics. Perplexity sends a referrer, so you can filter for `perplexity.ai` (and the Comet browser's referrer) in GA4 to see the traffic it drives. Pair that with a short list of your priority prompts run by hand every week or two, and you have a zero-cost baseline. The blind spot: when someone copies your URL out of an answer instead of clicking the citation, that visit lands in analytics as direct traffic, so referral numbers understate the channel. Treat the analytics figure as a floor, not the full picture. Free graders are the other option, and they are useful for a single read. HubSpot's AEO Grader and similar tools score your AI presence once at no cost, but a grade is a photo, not a security camera; it will not tell you that you lost a citation last week. Our own free layer sits here too: the AI Crawler Checker and [AI-readiness score](https://geotoolbox.ai/tools/ai-readiness) are free, and paid tracking — Perplexity included from Starter — opens with a 7-day free trial. The short version: the free tools get you a baseline and a reachability check, while scheduled Perplexity tracking with history and competitor share is what you pay for. ## What to Look for in a Perplexity Rank Tracker Once you are comparing paid tools, these are the things that separate a real tracker from a dashboard with a nice logo: 1. **It captures the actual cited URLs**, not just a yes/no that you were mentioned. You need to know which of your pages earned the citation to do anything about it. 2. **It separates citations from mentions.** A tool that lumps them together inflates your visibility and hides whether you are getting clicks or just name-checks. 3. **Its refresh rate matches the volatility.** Perplexity's results move, so daily to weekly sampling is the sane range. A tool that checks monthly is measuring noise. 4. **Perplexity is in the plan you are pricing.** This is the one people miss. Confirm the specific tier that includes Perplexity, because several tools ship it only on an upgrade. 5. **It shows reachability, or you check that separately.** A tracker that reports "not cited" without flagging that a crawler cannot reach you will send you rewriting content that was never the problem. 6. **Pricing is transparent and self-serve.** If you cannot see the price without a sales call, you cannot compare value. 7. **It helps you act.** Monitoring tells you what happened. The tools worth paying for point at what to change next. ## When You Don't Need a Perplexity Rank Tracker Here is the part most guides leave out: sometimes the answer is not to buy one. Skip a paid tracker for now if you publish fewer than about four substantive pieces a month, because you will not have enough new material for citations to move, and the tool mostly measures stagnation. Skip it if you are a purely local service business, since Perplexity tends to lean on maps and directories for those queries and a tracker adds little. Skip it if you run a product-only store with no editorial content, because there is not much for Perplexity to cite yet. And if your content budget is under about a thousand dollars a month, a two-hundred-dollar tracker is the wrong place to spend; a quarterly manual check plus a free reachability test covers you. We would rather tell you that than sell you a subscription you will not use. When the volume and stakes are there, geotoolbox and the tools above earn their keep. When they are not, the right tool is a spreadsheet and thirty free minutes a quarter. ## The Verdict There is no single best Perplexity rank tracker, and any list that hands you one is usually selling it. The right tracker covers the engines you care about, at a tier where Perplexity is actually included, and gives you something to act on. To make it concrete: - **Otterly** if budget comes first and you want cheap multi-engine coverage - **Rankability** if you need the exact URLs Perplexity cites, not just a score - **ZipTie** if you want fix recommendations, not only a dashboard - **Peec AI** if the priority is benchmarking against named competitors - **Nightwatch** if you want Perplexity citations sitting next to your Google rank tracking - **Geotoolbox** (ours) if you want Perplexity included at the entry price, in context with the seven other engines, plus the reachability check that rules out a blocked crawler first - an SEO-suite add-on (**Semrush**, **SE Ranking**) if you already pay for the suite; **Scrunch** or **Profound** if you are enterprise and can price the Perplexity tier But before any of that, do the free step first: confirm the AI crawlers can reach and read your key pages, since that is the one problem no amount of tracking solves. Run your site through our free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) to rule it out in a minute, and if you would rather see Perplexity next to the seven other engines than watch it alone, a [GEO Scan](https://geotoolbox.ai/features/geo-scan) shows where you are cited and where a competitor took your spot. Fix the reachability, then track what moves. ## Frequently Asked Questions ### Does Perplexity have rankings like Google? No. Perplexity does not order pages one through ten. It runs a live search, reads a handful of pages, and cites roughly three to five of them in a synthesized answer. A "rank tracker" for Perplexity really tracks whether and how often you land in that cited set, not a fixed position. ### Is there a free Perplexity rank tracker? The genuinely free method is your own analytics: filter for the `perplexity.ai` referrer in GA4 and run your priority prompts by hand every week or two. Free graders like HubSpot's give a one-off snapshot. Continuous, scheduled tracking with history and competitor share is what paid tools charge for. ### What's the difference between a citation and a mention in Perplexity? A citation is a clickable numbered source link, and it drives referral traffic. A mention is your brand named in the answer text with no link, which builds awareness but sends no clicks. A good tracker reports the two separately, because they are different outcomes. ### Why do I get different Perplexity results for the same prompt? Because Perplexity searches live and its answers are non-deterministic. The cited sources shift with timing, location, personalization, and model version, so [research on AI-search visibility](https://arxiv.org/abs/2603.08924) treats a single check as one sample, not a fixed truth. That is exactly why you sample on a schedule instead of trusting one lookup. ### If I rank on Google, will Perplexity cite my site? Not necessarily. Perplexity has to reach the page and lift a clean, attributable claim from it. A page that ranks well on Google but buries its answer can lose to a shorter, clearer page, and a page blocked from [PerplexityBot](https://geotoolbox.ai/glossary/perplexitybot) will not be cited at all. ### How often should I check my Perplexity visibility? Daily to weekly is the useful window, because results move but not by the hour. Monthly checks measure noise more than signal. Match the cadence to how often you publish and how competitive your space is. ## Sources - From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization - arXiv (Zhang et al.) - `arxiv.org/abs/2604.25707` - Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement - arXiv (Sielinski) - `arxiv.org/abs/2603.08924` - PerplexityBot and Perplexity-User crawler documentation - Perplexity - `docs.perplexity.ai/guides/bots` - How traffic from AI search converts - Seer Interactive - `seerinteractive.com/insights/case-study-6-learnings-about-how-traffic-from-chatgpt-converts` --- ## The Best Open-Source Tools for Claude Code (2026) > The open-source repos that make Claude Code better: MCP servers, subagent frameworks, cost trackers, and hooks, with the security caveats most roundups skip. - Canonical: https://geotoolbox.ai/blog/claude-code-open-source-tools - Published: 2026-07-11 · Updated: 2026-07-22 [Claude Code](https://geotoolbox.ai/blog/what-is-claude-code) is good the moment you install it. It gets genuinely powerful when you wrap it in the tooling the community has built around it. Most people run it like early autocomplete (type a request, take the answer, move on) and never touch the layer that turns it into a real engineering environment: memory, live docs, browser control, cost visibility, and reusable workflows. That layer is open source, mostly free, and installable in an afternoon. The most useful tools cluster into a few jobs: MCP servers that give Claude new capabilities (the official reference set, Context7, GitHub, Playwright), workflow and subagent frameworks (Superpowers, SuperClaude), and cost trackers (ccusage). This is the shortlist worth knowing, grouped by the job each tool does. Every repo here is real and actively maintained, and the star counts are current as of July 2026 (they move fast, so treat them as a rough signal of adoption, not a leaderboard). We also cover how to vet these tools before you hand them access to your machine. ## Enhance, Don't Replace: Two Different Questions Search for "open-source tools for Claude Code" and you get two different answers tangled together. One is tools that **replace** Claude Code: open-source CLI agents like OpenCode, Aider, and Cline that run the same job on a different (often model-agnostic) engine, up to and including a model you [run locally yourself](https://geotoolbox.ai/blog/run-llm-locally). The other is tools that **improve** the Claude Code you already use. This guide is about the second kind. Nothing below is an alternative to Claude Code. Every tool assumes you are running Claude Code and want it to do more, see more, or cost less. If you are shopping for a replacement instead, that is a separate question with a separate answer, and OpenCode is the usual starting point. ## What Makes a Tool Worth Adding A tool earns a spot in your setup when it clears three bars: it installs in a few minutes, it solves a job you hit repeatedly, and it is maintained (recent commits, real users, an issue tracker that gets answered). Star count is a proxy for the last one, not a reason on its own. The counter-intuitive part is knowing when to stop. More tools is not better. Every MCP server you connect adds its tools to the list Claude has to choose from on each turn, and past a point that longer menu actively degrades tool-selection accuracy: the model picks the wrong tool, or wastes context deciding. Independent write-ups keep landing on the same advice: install a focused set, start with three or four, and add more only when a real workflow demands it. A lean setup that Claude uses well beats a maximal one it fumbles. ![The Claude Code tool stack by job: discovery (awesome-claude-code lists), MCP servers (official servers, Context7, GitHub, Playwright, Serena), frameworks and subagents (Superpowers, SuperClaude, wshobson/agents, BMAD), cost and context (ccusage, ccstatusline, Usage-Monitor), and hooks, all layered on top of the Claude Code CLI, with a 'vet first' security note above the stack.](/blog/claude-code-open-source-tools/claude-code-tool-stack-by-job.png) The rest of this guide walks up that stack, one job at a time. ## Start Here: The Community Indexes Before installing anything specific, bookmark the two lists the whole ecosystem starts from. They save you from chasing dead repos and let you see what exists by category.
RepoStars (Jul 2026)What it is
hesreallyhim/awesome-claude-code~50kThe curated index of skills, hooks, slash commands, subagents, MCP servers, and workflows. The map everyone starts from.
VoltAgent/awesome-claude-code-subagents~23kA library of 100+ ready-made subagents (code reviewer, security auditor, architect, debugger) you drop into .claude/agents.
## MCP Servers That Give Claude Code New Senses MCP servers are usually the highest-leverage layer to add first, because each one hands Claude a capability it otherwise lacks and can call from inside the terminal: reading live docs, driving a browser, querying a database, searching your codebase by meaning instead of filename. Start with the official reference servers, then add the ones that match your work.
RepoStars (Jul 2026)What it adds
modelcontextprotocol/servers~88kThe official reference set: filesystem, git, memory, sequential-thinking, fetch. The canonical place to begin.
upstash/context7~59kLive, version-correct documentation on demand, so Claude stops inventing outdated APIs and wrong function signatures.
microsoft/playwright-mcp~35kBrowser automation. Claude navigates your app, fills forms, takes screenshots, and verifies UI flows end to end.
github/github-mcp-server~31kOfficial GitHub server. Read issues, open and review PRs, search code, inspect Actions. Especially useful when most of your work lives in GitHub.
oraios/serena~26kSemantic code retrieval and editing through a language server, so Claude works at the symbol level, not blind text search.
zilliztech/claude-context~12kIndexes your whole repo for semantic search, useful on monorepos and large legacy codebases.
firecrawl/firecrawl-mcp-server~7kTurns any web page into clean markdown for onboarding Claude to a new framework or feeding a RAG workflow.
Most install in one line. The official servers follow the pattern `claude mcp add filesystem -- npx -y @modelcontextprotocol/server-filesystem /path/to/root`, and the others document their own one-liner. If you only add two, make them the official servers and GitHub. Two jobs sit just outside that table but come up constantly. For databases, [crystaldba/postgres-mcp](https://github.com/crystaldba/postgres-mcp) lets Claude inspect a schema and run queries against Postgres (use a read-only connection). For live web search rather than scraping, [exa-labs/exa-mcp-server](https://github.com/exa-labs/exa-mcp-server) gives Claude a real search tool. Memory and step-by-step reasoning are worth having too, and both ship inside the official reference set above (the `memory` and `sequential-thinking` servers), so you get them for free with the canonical install. ## Skills, Subagents, and Orchestration Frameworks The next layer is workflow. Out of the box, Claude Code takes a request and runs. Frameworks give it structure: reusable skills, specialized subagents, and multi-step methods that turn "build this feature" into plan, then test, then implement, then verify.
RepoStars (Jul 2026)What it does
obra/superpowers~250kA skills framework and methodology: battle-tested skills for planning, TDD, debugging, and structured execution pipelines.
ruvnet/ruflo~64kMulti-agent orchestration (formerly claude-flow). Coordinates swarms of agents working in parallel.
bmad-code-org/BMAD-METHOD~50kAn artifact-driven agile method: analyst, PM, architect, and developer roles handing work down a defined pipeline.
wshobson/agents~38kA marketplace of agentic plugins that works across Claude Code and other harnesses.
SuperClaude-Org/SuperClaude_Framework~24kA configuration framework that adds specialized commands and cognitive personas to Claude Code.
These split into camps. Superpowers and SuperClaude sharpen how a single Claude Code session works. Ruflo and BMAD are for multi-agent orchestration, running several agents against one problem. Pick based on whether your bottleneck is the quality of one session or the coordination of many. Avoid running two orchestration frameworks at once: they fight over the same files and conventions. ## Keep Token Costs and Context Under Control Claude Code's running cost comes from a quiet source: context accumulation. Every file it reads and every tool result stays in the window and gets re-billed on the next turn, so a long session becomes a token furnace you cannot see. The tooling here makes that spend and that context visible before it hurts.
RepoStars (Jul 2026)What it does
ccusage/ccusage~17kOne command (npx ccusage) reads your local logs and reports token use and cost by session, day, and model.
sirmalloc/ccstatusline~12kA customizable status line showing live model, context usage, and cost right in the CLI.
Maciek-roboblog/Claude-Code-Usage-Monitor~8kA real-time monitor with burn-rate predictions and warnings before you hit a limit.
Seeing your context fill up is what lets you act: compact, or start a fresh session, before quality drops. If token spend is your main pain, that is a topic in itself: our guide to [reducing Claude Code's token costs](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) goes deeper, and [how the context window actually works](https://geotoolbox.ai/blog/claude-code-context-window) explains why sessions decay. Hooks are the other customization layer worth knowing. They let you run your own scripts on Claude Code's events: a linter on every file write, a notification when a long task finishes, or a guard that blocks a risky command before it runs. [disler/claude-code-hooks-mastery](https://github.com/disler/claude-code-hooks-mastery) is the reference to read first, and [carlrannaberg/claudekit](https://github.com/carlrannaberg/claudekit) bundles ready-made hooks and commands you can adopt wholesale. ## Before You Install Anything: Vet Your MCP Servers This is the part most roundups skip, and it matters more than any single tool. An MCP server is not a passive plugin. It runs with your permissions and can read files, execute commands, and reach the network on Claude's behalf. A malicious or careless one is a security hole you opened yourself. The risk is not hypothetical. Security researchers found critical flaws in Anthropic's own Git MCP server, [tracked as CVE-2025-68145, CVE-2025-68143, and CVE-2025-68144](https://www.techradar.com/pro/security/anthropics-official-git-mcp-server-had-some-worrying-security-flaws-this-is-what-happened-next): a path-validation bypass, an unrestricted init, and an argument injection. Chained with the Filesystem server, they could reach remote code execution through a prompt injection. Anthropic patched the flaws in December 2025, the details were disclosed in January 2026, and no exploitation in the wild was reported. The lesson stands anyway: even official, well-reviewed servers ship real vulnerabilities. Unvetted community packages are a bigger unknown, and the `--dangerously-skip-permissions` flag, which lets Claude run commands without stopping to ask, turns any bad instruction straight into action. Neither is a reason to avoid the ecosystem, just a reason to treat access as something you grant deliberately. Treat every MCP server as a privileged extension. Install only from trusted, well-maintained vendors. Use read-only credentials wherever possible, scope filesystem and shell access as tightly as you can, and never grant unrestricted shell access to a server you have not read. Reserve `--dangerously-skip-permissions` for throwaway sandboxes, never a machine with anything you would miss. ## A Sensible Starter Stack If you are setting up a new machine today, resist the urge to install everything at once. A focused stack you understand beats a sprawling one Claude cannot navigate. Start with five tools, only three of them MCP servers: the official reference servers (filesystem, git, memory), the GitHub server, and Context7 for live docs, plus ccusage for cost visibility and one workflow framework (Superpowers is the safe default). Add a browser or codebase-search server only when a real task needs it. ## Where This Fits With AI Visibility Tooling like this makes Claude Code a stronger builder. It is worth flipping that around: once you are wiring agents into your own products, how do the AI engines describe *your* brand when someone asks about your space? That is the job we work on at geotoolbox. If you are already building with agents, it is worth knowing whether the models can reach your site and what they say about you. You can [check how your brand shows up across the AI engines for free](https://geotoolbox.ai/tools/ai-readiness), and if you want the wider picture, our roundup of [generative engine optimization tools](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools) covers the measurement side. For the model itself, start with [what Claude AI is](https://geotoolbox.ai/blog/what-is-claude-ai). ## Frequently Asked Questions ### What is the difference between a plugin, a skill, an MCP server, and a subagent? They sit at different layers. A **skill** is a reusable instruction set Claude Code loads for a task. A **subagent** is a separate Claude instance with its own focus (a reviewer, a debugger) that the main session delegates to. An **MCP server** is an external tool Claude can call to reach outside itself: a database, a browser, live docs. **Plugin** is the loose umbrella term for anything you add on top. Most setups mix all four. ### Are these open-source tools free? The repos listed here are open source and free to run. A few lean on a hosted backend (Context7's server is open source, for instance, but its documentation index is a hosted service), and in every case you still pay for the model usage underneath: Claude Code bills tokens whether or not these tools are installed. Some, like the cost trackers, exist specifically to keep that bill visible. ### Which MCP servers should I install first? The official reference servers (filesystem, git, memory) and the GitHub server cover most day-to-day work. Add Context7 so Claude checks current docs instead of guessing at function signatures. That set handles a large share of real tasks before you need anything specialized. ### Is it safe to install community MCP servers? Only with care. An MCP server runs with your permissions, so an untrusted one is a genuine risk, as the Git MCP vulnerabilities showed even for official code. Install from maintained, reputable projects, use read-only credentials, and scope access tightly. When in doubt, read the source before you connect it. ### What is the best open-source alternative to Claude Code? That is a different question from this guide, which is about tools that enhance Claude Code rather than replace it. If you specifically want an open-source CLI agent instead of Claude Code, OpenCode, Aider, and Cline are the names that come up most, all model-agnostic and free to run against your own API keys. ### How many MCP servers is too many? There is no hard number, but more is not better. Each connected server adds tools to the menu Claude picks from every turn, and a longer menu lowers its accuracy at choosing the right one. Start with three or four, and add another only when a workflow clearly needs it. ## Sources - hesreallyhim/awesome-claude-code - the community index of Claude Code skills, hooks, commands, and MCP servers - `github.com/hesreallyhim/awesome-claude-code` - modelcontextprotocol/servers - the official reference MCP servers - `github.com/modelcontextprotocol/servers` - upstash/context7, oraios/serena, microsoft/playwright-mcp, github/github-mcp-server - the most-adopted MCP servers referenced above - `github.com/upstash/context7` - `github.com/oraios/serena` - `github.com/microsoft/playwright-mcp` - `github.com/github/github-mcp-server` - obra/superpowers, ruvnet/ruflo, ccusage/ccusage - the workflow, orchestration, and cost tools referenced above - `github.com/obra/superpowers` - `github.com/ruvnet/ruflo` - `github.com/ccusage/ccusage` - The best GitHub repos for Claude Code, June 2026 - Markus Stöger - starter-stack picks and the case against over-installing - `markusstoeger.com/en/masterai/best-claude-code-repos-2026` - Best Claude Code MCP Servers in 2026: Setup and Top 10 - Totalum - install commands and MCP security rules - `totalum.app/blog/claude-code-mcp-servers-2026` - Anthropic's official Git MCP server had some worrying security flaws - TechRadar, January 2026 - the Git MCP CVEs and fix - `techradar.com/pro/security/anthropics-official-git-mcp-server-had-some-worrying-security-flaws-this-is-what-happened-next` --- ## DeepSeek Pricing in 2026: Free App, Cheap API, Is It Worth It? > DeepSeek pricing, current to August 2026: the free app, V4 Flash and V4 Pro API costs, cache-hit vs cache-miss billing, and how it compares to ChatGPT and Claude. - Canonical: https://geotoolbox.ai/blog/deepseek-pricing - Published: 2026-07-11 · Updated: 2026-08-18 The DeepSeek app is free. That is the whole answer for most people who type "DeepSeek pricing" into Google. You pay nothing to chat with it on the web or in the app, and there is no Plus or Pro subscription to upsell you. The money question only starts when you move to the API. And there the news is good: DeepSeek is still close to the cheapest frontier-class pricing you can buy, with V4 Flash starting at $0.22 per million input tokens off-peak and V4 Pro at $0.66 input / $1.98 output off-peak, as of August 2026. The catch, and it is a new one: DeepSeek's long-threatened "significant" price increase landed on August 16, 2026, replacing the old flat rate with a peak/off-peak split that roughly doubles the price during seven hours a day. Its closest open-weight rival on price is Moonshot's Kimi, though the arrival of K3 moved that lineup upmarket; see our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) breakdown and, for the consumer plans, our [Kimi pricing](https://geotoolbox.ai/blog/kimi-pricing) guide. The other catch is that many of the pricing guides you will find still quote the April 2026 launch price of $1.74 and $3.48 for V4 Pro, or the July flat rate of $0.435 and $0.87, both now stale. Below is every current price, reconciled and dated, the two things that move your bill most, and whether it is worth paying for. ## How Much Does DeepSeek Cost? Free App vs Paid API DeepSeek is sold as two different things, and almost all the confusion comes from mixing them up. The **consumer product** is the DeepSeek chat app on the web and phone. It is free, has no subscription tier, and is what most searchers want. The **API** is a separate, developer-facing service you pay for per token, used to wire DeepSeek into your own apps, agents, or coding tools. Free chat access does not include free API usage.
What you useWhat it costsWho it is for
DeepSeek app (web + mobile)Free, no subscriptionAnyone chatting with DeepSeek directly
DeepSeek API - V4 FlashFrom $0.22 / $0.66 per 1M tokens off-peak (up to $0.44 / $1.32 at peak hours)High-volume apps, chatbots, extraction, most coding
DeepSeek API - V4 Pro$0.66 / $1.98 per 1M tokens off-peak (up to $1.32 / $3.96 at peak hours)Hard reasoning, complex code, agentic work
If you just want to use DeepSeek, stop at the free app and skip the rest of this page. If you are building on it, the sections below are where the real cost lives, and it is not the sticker price. For the bigger picture of what the models are and where DeepSeek came from, our guide to [what DeepSeek is](https://geotoolbox.ai/blog/what-is-deepseek) covers the history and the open-weights story in plain terms. ## Is DeepSeek Free? What the App Actually Gives You Yes. The DeepSeek app is free, with no paid consumer plan at all. Sign in at chat.deepseek.com or the mobile app and you get the full chat experience, including the V4 model, web search, and file uploads, at no charge. Unlike [ChatGPT Plus](https://geotoolbox.ai/blog/chatgpt-pricing) or [Claude Pro](https://geotoolbox.ai/blog/claude-pricing) at around $20 a month, DeepSeek does not sell a fixed monthly subscription for individual use. The reason is straightforward: the consumer chat is already free, and the API is prepaid per token, so there is nothing to bundle into a monthly plan. The one limit worth knowing is fair-use throttling. During heavy traffic you may see "Server Busy" messages and slower responses, which is DeepSeek managing load rather than a paywall. There is no per-day message quota the way some rivals cap their free tiers. The other thing to keep in mind: the app is text-only. DeepSeek does not generate images, video, or audio, so if you came looking for a free image generator, this is not it. ## DeepSeek API Pricing: V4 Flash and V4 Pro If you are building on DeepSeek, you pay per [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), the chunks of text the model reads and writes, billed separately for input and output, and, since August 16, 2026, separately again for peak and off-peak hours. There are two current models, and the lineup is deliberately simple. Here are the rates from [DeepSeek's official API docs](https://api-docs.deepseek.com/quick_start/pricing).
ModelInput, cache miss (per 1M)Input, cache hit (per 1M)Output (per 1M)Context
deepseek-v4-flash — off-peak$0.22$0.007$0.661M
deepseek-v4-flash — peak$0.44$0.014$1.321M
deepseek-v4-pro — off-peak$0.66$0.022$1.981M
deepseek-v4-pro — peak$1.32$0.044$3.961M
Peak hours are 01:00-04:00 and 06:00-10:00 UTC; every other hour is off-peak. Off-peak rates are exactly half of peak across every column, so the fastest way to estimate your bill is to price it at off-peak and then check what share of your traffic falls in those two UTC windows. Both models share a one-million-token context window and up to 384K tokens of output, with no separate long-context surcharge, so a huge prompt bills at the same per-token rate as a small one at whatever time of day it runs. V4 Flash is the default workhorse: chat, extraction, classification, summarization, and most coding. V4 Pro is the stronger model for hard reasoning, multi-step analysis, and demanding code, and it costs roughly three times Flash on cache-miss input and output, at either peak or off-peak rates. Two naming details trip people up. First, if you have seen `deepseek-chat` and `deepseek-reasoner` in older tutorials, those are legacy aliases of V4 Flash (non-thinking and thinking mode), and DeepSeek retired both names on July 24, 2026, so new code must call `deepseek-v4-flash` and `deepseek-v4-pro` directly. Second, on the hosted API the `deepseek-reasoner` alias mapped to V4 Flash's thinking mode, not to Pro; R1 was DeepSeek's earlier standalone reasoning model, and the docs point new work to V4 Flash and V4 Pro. The cache-hit and cache-miss columns are not a typo. That roughly thirty-fold gap between $0.22 and $0.007 on off-peak Flash input is the single most important thing to understand about DeepSeek pricing, and it is why your real bill rarely matches the number you first read. The next two sections cover why. ## Cache-Hit vs Cache-Miss: Why the Sticker Price Isn't Your Bill DeepSeek charges two different input prices, and which one you get depends on context caching. Caching is automatic and on by default. When the beginning of your request matches something DeepSeek processed recently (a reused system prompt, a shared set of instructions, a document you keep sending), those matching tokens bill at the cache-hit rate instead of the full input rate. On V4 Flash that drops input from $0.22 to $0.007 per million tokens off-peak (or $0.44 to $0.014 at peak), a 97% reduction either way. On V4 Pro the cache-hit rate cuts input by roughly the same margin, an even bigger discount in absolute dollars. There is no separate cache-write charge and no hourly storage fee, so caching is pure savings. The condition is that cache hits need an exact matching prefix. Put your static content, the system prompt and instructions, at the very start of the request, and your variable content at the end. A small change near the front breaks the match and sends those tokens back to the full cache-miss rate. Caching also works on a best-effort basis, so rather than assuming every repeated prompt hits the cache, track the `prompt_cache_hit_tokens` and `prompt_cache_miss_tokens` fields DeepSeek returns with each response and tune from there. This is why quoting "DeepSeek Flash is $0.22" is only half true. A workload that reuses a long system prompt across thousands of calls pays far less than that on input; a workload that sends a fresh prompt every time pays the full rate, and whether that request lands in a peak or off-peak UTC window moves it again. Your architecture and your traffic's time zone, not the price list, decide which bill you get. ## Why Your DeepSeek Bill Runs Higher Than the Off-Peak Sticker Price The other surprises are on the output side and the clock. Both V4 models run in thinking mode by default. Before it answers, the model generates internal reasoning tokens, and those tokens bill at the output rate even though they are not part of the final answer. So a request that returns a two-line answer can quietly produce thousands of billed reasoning tokens first. If your bill is higher than the sticker price suggested, thinking mode left on for simple tasks is the most common reason. Turning it off for routine work (classification, extraction, short replies) cuts output tokens directly. Output is also the expensive side of the ledger. On both models it costs three times the off-peak cache-miss input rate, and combined with a 384K maximum output and default thinking, runaway responses are the main way DeepSeek bills get "weird." The peak-hour window adds a third lever: 01:00-04:00 and 06:00-10:00 UTC bill everything, input and output alike, at double the off-peak rate, so a workload that happens to run heavy during those seven hours a day pays noticeably more than the same volume shifted a few hours earlier or later. Cap your output length, ask for concise or structured responses, batch what you can outside peak hours, and reserve the long generations for tasks that genuinely need them. ## What DeepSeek Costs Per Month, in Real Dollars Per-token rates are hard to feel, so here is what typical workloads actually cost on V4 Flash at off-peak rates, the model most production traffic should use. Peak-hour traffic runs roughly double every figure here.
WorkloadRough monthly cost (off-peak)Model
Customer support chatbot, ~1,000 conversations~$3.50 (under $2 with caching)V4 Flash
Same chatbot on the stronger model~$11.00V4 Pro
Summarizing 100 PDFs~$0.45V4 Flash
30 articles of content generation~$0.08V4 Flash
The pattern holds at scale. Light personal use tends to land around $2 to $8 a month off-peak, small production apps around $8 to $40, and heavier production workloads in the low hundreds, more if traffic clusters in the 01:00-04:00 or 06:00-10:00 UTC peak windows. Those are Flash numbers; moving the same volume to V4 Pro multiplies them by roughly three. New API accounts have often been given a small promotional grant of free tokens to test with, though the amount and expiry change over time, so check your balance on the platform after signing up rather than trusting a fixed figure from a guide. ## How DeepSeek Billing Works, and How to Start The API runs on a prepaid balance, not a subscription or a monthly invoice. You top up a credit balance, and each call draws it down. There is no seat fee and no auto-renewing plan, so your spend is capped at whatever you loaded. The gotcha that catches people: if the balance hits zero, requests start failing with a `402 Insufficient Balance` error even though your API key is still valid and the free app keeps working normally. Topping the balance back up clears it immediately, and unused credit does not expire. It is worth setting a low-balance alert if anything production depends on the API. One thing that makes DeepSeek easy to adopt is that the API is OpenAI-compatible and Anthropic-compatible. In most existing code you point the `base_url` at DeepSeek and swap the key rather than rewriting anything, which is a big part of why teams move workloads onto it to cut costs. ## Is DeepSeek Cheaper Than ChatGPT, Claude, and Gemini? On the API, yes, and by a wide margin. That price gap is much of why DeepSeek matters commercially.
ModelInput (per 1M)Output (per 1M)
DeepSeek V4 Flash (off-peak)$0.22$0.66
DeepSeek V4 Pro (off-peak)$0.66$1.98
Grok 4.3$1.25$2.50
Gemini 3.1 Pro~$2.00~$12.00
Claude Opus 5~$5.00~$25.00
GPT-5.6 Sol~$5.00~$30.00
Against a top-tier model like Claude Opus 5, V4 Pro is now roughly 7.5 times cheaper on input and nearly 13 times cheaper on output at off-peak rates, narrower than before the August 16 price change but still a wide gap; at peak hours that shrinks further to roughly 3.8 and 6.3 times. V4 Flash widens the off-peak gap further. These are flagship tiers, which is the honest frame: each rival also runs a budget model, and a couple dip below DeepSeek Flash on off-peak input (Google's Gemini 2.5 Flash-Lite is $0.10), while every Western provider offers a 50% batch discount on bulk jobs that DeepSeek has no answer to. Grok 4.3 in the table is xAI's value model rather than its flagship. Against the top tier of each rival, though, DeepSeek still sits in a cheaper bracket, just a less dramatic one than it did a week ago.
![Bar chart drawn to scale comparing API output price per million tokens in August 2026, at DeepSeek's off-peak rate: DeepSeek V4 Flash at $0.66 is far smaller than every rival shown, V4 Pro at $1.98 sits just under Grok 4.3's $2.50 and is far smaller than Gemini 3.1 Pro at $12, Claude Opus 5 at $25, and GPT-5.6 Sol at $30.](/blog/deepseek-pricing/deepseek-vs-frontier-output-cost.png)
Output tokens per 1M, to scale against a $30 axis, at DeepSeek's off-peak rate. Peak-hour DeepSeek pricing (01:00-04:00 and 06:00-10:00 UTC) is double these bars.
A headline "ten times cheaper" is not your real savings, though. True cost depends on how much output you generate, whether thinking mode is on, whether your traffic falls in DeepSeek's peak window, and how often you hit the cache, and DeepSeek's default thinking mode can erode part of the advantage on output-heavy work. It is also worth weighing the non-price factors, covered below, before switching a production system on price alone. For a feature-level view rather than just the token math, our comparison of [Chinese AI models](https://geotoolbox.ai/blog/chinese-ai-models-compared) puts DeepSeek next to its closest rivals. ## The Stale-Price Trap: Three Different "Current" Prices Are Floating Around Here is what most DeepSeek pricing guides get wrong right now: there have been three distinct V4 Pro prices in under five months, and a lot of the web still shows the first two. V4 Pro launched in April 2026 with a list price of $1.74 per million input tokens and $3.48 output. DeepSeek cut that by about 75% for a flat rate of $0.435 and $0.87 that held through most of the summer. Then, on August 16, 2026, DeepSeek made good on the "significant increase" it had been warning about since July: the flat rate was replaced with the peak/off-peak split above, which puts V4 Pro at $0.66 / $1.98 off-peak and $1.32 / $3.96 at peak hours. A lot of guides, aggregators, and even calculators still quote one of the two earlier numbers. It gets muddier because third-party hosts that resell DeepSeek, listed on marketplaces like [OpenRouter](https://openrouter.ai/deepseek/deepseek-v4-pro), set their own rates, so a figure you find there may not match DeepSeek's own price at all. If you build a budget on the original $1.74 launch price, your cost model now overstates off-peak input cost by roughly 2.6 times and off-peak output cost by roughly 1.8 times, and at peak hours the old $3.48 output figure actually undercounts, since the current peak output rate ($3.96) runs about 14% above it. Even the July flat rate undercounts what you will pay today: roughly 50% low on off-peak input (up to about 200% low at peak input), and 128% low on off-peak output (up to roughly 355% low at peak output). When you read a DeepSeek price, check the date, and treat any V4 Pro figure that isn't peak/off-peak-aware as out of date. One older detail still confuses people: DeepSeek's earlier off-peak discount, the V3-era window that knocked 50 to 75% off during quiet hours, disappeared when V4 launched with flat all-day rates, then reappeared in this new form on August 16 with different hours and a smaller discount (peak is now exactly double off-peak, not the old 2-4x V3 spread). Don't assume the current peak/off-peak windows match the ones from the V3 era; the UTC hours cited above (01:00-04:00 and 06:00-10:00) are the current ones as of this pricing page. ## Which DeepSeek Option Should You Actually Use? Short answer to the title question: for API work, yes, it is worth it wherever text-only output and China hosting are acceptable. The rest is picking the right tier. **Use the free app** if you want to chat with DeepSeek, run research, or draft text. It costs nothing, needs no subscription, and covers what the large majority of searchers are after. **Use V4 Flash on the API** for almost everything you build: chatbots, extraction, classification, summarization, and most coding. It is cheap enough that at typical volumes the bill is a rounding error, and DeepSeek positions it as matching Pro on simpler agent tasks. **Reserve V4 Pro** for the work that genuinely needs it, hard reasoning, complex multi-step analysis, and demanding code, where the quality gain is worth roughly three times the token cost. **Look elsewhere** when price is not the only constraint. DeepSeek is text-only, so it will not cover image, video, or audio generation. Its models are hosted in China, which matters if you have data-residency, GDPR, or industry compliance requirements, and the pricing docs say nothing about data-processing agreements. For a regulated business, the cheapest token is not automatically the right one. ## Why DeepSeek's Price Matters for Your Brand DeepSeek being this cheap is exactly why it shows up everywhere. It is the default model builders reach for when they want frontier-ish quality at throwaway cost, which means it is quietly powering the chatbots, agents, and search tools that describe your category to buyers. The more useful question, then: when someone asks an AI assistant what to buy or who to hire, does it mention you? In our tracking of AI answers, assistants routinely name specific vendors in response to buying questions, and the buyer rarely sees why one brand surfaced and another did not. None of that shows up on a pricing page. Our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) maps which sources the AI engines cite when they answer, so you can find the ones shaping the recommendation and get into them. The method is in our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). DeepSeek has changed its API pricing twice already this year, once down and now once up, so confirm the current figure on the official page before you build a budget, and once you know what it costs, the next question is what it tells people about you. ## Frequently Asked Questions ### Is the DeepSeek app free, or is there a subscription? The app is free, and there is no paid consumer subscription. You get the full chat experience, including the V4 model, web search, and file uploads, at no charge on the web and mobile app. The only thing you pay for is the API, which is billed per token for developers building their own software. ### Why is my DeepSeek API bill higher than the advertised price? Usually because of thinking mode, which is on by default and bills its reasoning tokens at the output rate, so a short answer can cost far more than it looks. The other reason is caching: the low input price only applies to cache hits, and a prompt that changes every time pays the higher cache-miss rate. Disable thinking for routine tasks and keep your prompt prefixes identical to control both. ### What is cache-hit vs cache-miss pricing? DeepSeek charges two input prices. When the start of your request matches content it processed recently, those tokens bill at the cache-hit rate, which is about 97% cheaper on V4 Flash. When the input is new, it bills at the full cache-miss rate. Caching is automatic and free to use, but it only triggers on an exact matching prefix. ### Is DeepSeek really cheaper than OpenAI, Claude, and Gemini? On the API, yes, by a wide margin, though a narrower one since DeepSeek's August 16, 2026 price change: V4 Pro is roughly 7.5 times cheaper than Claude Opus 5 on input and nearly 13 times cheaper on output at off-peak rates (about 3.8x and 6.3x at peak hours), and V4 Flash is cheaper still. Your real savings depend on how much output you generate, how often you hit the cache, and whether your traffic falls in DeepSeek's peak window, but on raw token price DeepSeek still undercuts every Western flagship, though a few rivals' budget tiers get close on cheap, batchable work. ### What happened to R1, V3, and deepseek-chat? V4 Flash and V4 Pro are the current models. The older `deepseek-chat` and `deepseek-reasoner` names are legacy aliases of V4 Flash and were retired on July 24, 2026, so new code must call the V4 model names directly. R1 was the earlier standalone reasoning model, now succeeded by the V4 line's built-in thinking mode, and V3 has likewise been superseded by V4. For the full model picture, see our [DeepSeek V4 guide](https://geotoolbox.ai/blog/deepseek-v4). ### How much does the DeepSeek API cost per month? For most workloads on V4 Flash, still not much, though it moved up on August 16, 2026 when peak/off-peak pricing replaced the old flat rate: a support chatbot handling around 1,000 conversations runs about $3.50 a month at off-peak rates, and lighter jobs cost pennies. Light personal use lands around $2 to $8 a month off-peak, small production apps around $8 to $40, and the same volume on V4 Pro costs roughly three times as much; traffic that clusters in the 01:00-04:00 or 06:00-10:00 UTC peak window runs about double. ## Sources - Models & Pricing - DeepSeek API Docs - `api-docs.deepseek.com/quick_start/pricing` - DeepSeek Platform (API access) - `platform.deepseek.com` - Claude plans and pricing - Anthropic - `claude.com/pricing` - Google AI Pro and Ultra subscriptions - Gemini - `gemini.google/subscriptions` - OpenAI API pricing - OpenAI - `developers.openai.com/api/docs/pricing` - DeepSeek V4 Pro - API pricing and benchmarks - OpenRouter - `openrouter.ai/deepseek/deepseek-v4-pro` --- ## Grok 4.5: Specs, Benchmarks, Pricing & the Honest Verdict (2026) > Grok 4.5 is xAI's cheaper, faster 'Opus-class' coding model, out July 8, 2026. An honest look at specs, real benchmarks, true cost, EU access, and AI search. - Canonical: https://geotoolbox.ai/blog/grok-4-5 - Published: 2026-07-11 · Updated: 2026-08-13 Grok 4.5 is xAI's cheaper, faster answer to the frontier: a coding-focused model that Elon Musk calls "Opus-class," shipped on July 8, 2026. The reality is narrower than the headline. It sits a notch below the top models on raw intelligence, but it does frontier-adjacent work for a fraction of the cost, and it is fast. This covers what Grok 4.5 actually is, what it really costs, and why a new AI model matters even if you never write a line of code. (xAI has since made [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) its flagship, released August 12, 2026 at the same $2/$6 price and 500K context; this page covers 4.5 specifically.) ## Is Grok 4.5 Out? What xAI Just Launched Yes. Grok 4.5 went public on July 8, 2026, released by SpaceXAI, the entity xAI now operates under. It is a model built for coding, agentic tasks, and knowledge work, not a chatbot upgrade, and it shipped as the default model inside xAI's coding agent, Grok Build. The headline framing came from Elon Musk, who called it an "Opus-class model, but faster, more token-efficient, and lower cost." That claim is doing a lot of work. If you have used [xAI's Grok](https://geotoolbox.ai/blog/what-is-grok) before, the important shift is that 4.5 is aimed at people who ship software and do professional work, not at people asking a chatbot for trivia. One catch shaped the launch: Grok 4.5 was not available in the European Union on day one. According to [xAI's announcement](https://x.ai/news/grok-4-5), every EU market was excluded at the July 8 launch, with availability "expected in mid-July" while xAI completed the EU AI Act evaluations that apply to a model classified as carrying systemic risk. It hit that window: xAI confirmed on July 16, 2026 that Grok 4.5 is now fully available across Europe, so EU users can reach it through SpaceXAI's products and the API. ## What's New in Grok 4.5 The defining detail is who helped train it. Grok 4.5 was trained alongside Cursor, the AI coding editor, using real developer-session data. [Cursor's launch note](https://cursor.com/blog/grok-4-5) describes it as the first model built for more than software engineering, tuned on long-running, in-codebase agent workflows rather than static code snippets. That is the source of most of its coding strength, and also of one of its credibility questions. The training run itself was large. xAI says Grok 4.5 was trained across tens of thousands of NVIDIA GB300 GPUs, with reinforcement learning spanning hundreds of thousands of multi-step engineering tasks and heavy data filtering and deduplication. The stated goal was per-token intelligence, getting more out of each token rather than simply generating more of them. You will see claims that Grok 4.5 runs on a new "V9" foundation with roughly 1.5 trillion parameters, a jump from Grok 4.3's smaller base. Treat those numbers as unconfirmed. They come from secondary coverage and leaks, not from xAI's own model card, which publishes no parameter count or architecture. If parameter counts matter to your decision, note that this one is unverified. For context on why "bigger" does not automatically mean "smarter," a [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) design can carry a huge parameter count while only activating a fraction per request. On the product side, Grok 4.5 is the default model in Grok Build, xAI's terminal coding agent that can run up to eight parallel sub-agents. It sits in the same release cadence that produced [the still-unreleased Grok 5](https://geotoolbox.ai/blog/grok-5): ship narrower, capability-specific models fast rather than wait on one general flagship. ## Grok 4.5 Benchmarks: What the Numbers Show This is where the marketing and the numbers diverge. xAI's prose says Grok 4.5 exceeds comparable leading models. Its own benchmark chart says something more mixed. Across four of the coding benchmarks xAI published, Claude Fable 5, Anthropic's most capable model, posts the top score on all four, and Grok 4.5 only stays close on one of them.
Coding benchmarkGrok 4.5Fable 5 (max)GPT-5.5 (xhigh)Opus 4.8 (max)
DeepSWE 1.0 (pass@1)62.0%66.1%64.31%55.75%
DeepSWE 1.1 (mini-swe-agent)53%70%67%59%
Terminal Bench 2.183.3%84.3%83.4%78.9%
SWE-Bench Pro (resolve rate)64.7%80.4%58.6%69.2%
The picture is competitive, not dominant. Grok 4.5 beats GPT-5.5 on SWE-Bench Pro but trails Opus 4.8 and Fable 5 there, and it is closest to the leader on Terminal Bench 2.1. These are also xAI's own figures ([reported by MarkTechPost](https://www.marktechpost.com/2026/07/08/spacexai-releases-grok-4-5/) from the launch chart), with competitor numbers pulled from their system cards and leaderboards, so treat them as a vendor's best case and verify on your own tasks. Independent testing lands in the same place. On the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/models/grok-4-5), Grok 4.5 scores 54, which puts it behind the frontier leaders rather than level with them: as of July 25, 2026 that is thirteenth of 190 models, behind Claude Opus 5, Fable 5, GPT-5.6 Sol and Opus 4.8. The same testing flags a reliability caveat: by Artificial Analysis's measures its hallucination tendency looks higher than its predecessor's, even as its factual knowledge improved. A model that knows more but states wrong answers with more confidence is exactly the failure mode you do not want in legal, financial, or client-facing work, so verify anything high-stakes rather than trusting a confident reply. Musk pitched Grok 4.5 as "Opus-class," but his more precise comparison was narrower: "roughly comparable to Opus 4.7, but much faster," referencing the previous Opus generation rather than the then-current Opus 4.8 (Anthropic has since moved the tier on again to [Opus 5](https://geotoolbox.ai/blog/claude-opus-5)). That gap between his own two framings tells you most of what you need to know. The coding numbers also deserve extra scrutiny for a specific reason: the model was co-trained with Cursor on Cursor's own workflows, and Cursor left its own CursorBench test out of the launch comparison after flagging a training overlap. That does not taint the other benchmarks, but it is a reason to run your own trials before trusting any leaderboard. If you want the fuller side-by-side, our [Grok vs Claude comparison](https://geotoolbox.ai/blog/grok-vs-claude) breaks down where each model actually pulls ahead.
![Grok 4.5 leans cheaper, faster, and more token-efficient; frontier models lead on benchmarks and accuracy.](/blog/grok-4-5/grok-4-5-vs-frontier-scorecard.png)
Where Grok 4.5 leans against the frontier on xAI's launch-chart benchmarks: it wins on cost and efficiency, the top models win on capability.
## Grok 4.5 Pricing: The $2/$6 Headline and the Real Cost Price is where Grok 4.5 makes its actual case, with one catch that the headline rate hides. It costs $2 per million input tokens and $6 per million output tokens, with cached input at $0.30 per million and a 500,000-token context window, per [xAI's model docs](https://docs.x.ai/developers/models/grok-4.5). Against the current frontier, that is more than 60% cheaper than the Opus tier and GPT-5.5 on prompts under 200,000 tokens. The catch is a price cliff at 200,000 tokens. xAI publishes a tiered table, not a flat rate: past that point input goes to $4, cached input to $0.60, and output to $12. And it is a cliff rather than a marginal tier, because in xAI's own words, "requests whose prompt reaches 200k tokens are billed at the higher rate for all tokens in the request." A 201,000-token request costs roughly double a 199,000-token one. On a model sold on its 500,000-token window, the advertised $2/$6 only holds across the bottom 40% of that window, which is worth knowing before you design a long-context workload around it. One quiet trade-off worth noting: that 500,000-token window is smaller than the multi-million-token windows some earlier Grok variants advertised, so if you were feeding the model very large codebases or documents, 4.5 is a step down on context.
ModelInput / 1MOutput / 1MContext
Grok 4.5$2 ($4 ≥200K)$6 ($12 ≥200K)500K
Claude Opus 5$5$251M
GPT-5.5$5$30Unconfirmed
GPT-5.6 Luna (cheap tier)$0.20$1.201.05M
The headline is real, but it is not your final bill. xAI's built-in API tools (web search, X search, code execution, document search) carry separate per-call fees, so a tool-heavy agent run costs more than the token math alone. Budget from a real workload, not just the headline rate. For most buyers the comparison that matters is against the model they already pay for. Grok 4.5 is confirmed on OpenRouter at the same $2/$6 and 500K context, [routed through the gateway](https://openrouter.ai/x-ai/grok-4.5) if you do not want to sign up for xAI directly. If you are weighing it against the OpenAI side, the new [GPT-5.6 family](https://geotoolbox.ai/blog/gpt-5-6) still undercuts it only at the Luna tier, but the July 30, 2026 price cut widened that gap sharply: Luna dropped 80% to $0.20/$1.20, so it is now a tenth of Grok 4.5's input price and a fifth of its output price, on a 1.05M window rather than 500K. Terra, the tier above, matches Grok's $2 input but doubles it on output. Against consumer Grok plans, our [Grok pricing breakdown](https://geotoolbox.ai/blog/grok-pricing) covers SuperGrok and the free tier. ## Speed and Token Efficiency: Where Grok 4.5 Actually Wins If the benchmarks are mixed and the price is a win, efficiency is the knockout. Grok 4.5 is served at roughly 80 tokens per second, quick for a model in this capability tier, and xAI claims about twice the token efficiency of comparable leading models overall, solving tasks in under half the steps. The efficiency number is the one worth internalizing. On SWE-Bench Pro, Grok 4.5 resolved tasks using about 15,954 output tokens on average, against roughly 67,020 for Opus 4.8, a gap of about 4.2x fewer output tokens per task ([xAI's own figure](https://x.ai/news/grok-4-5)). Output tokens are the expensive half of the bill, so that efficiency stacks on top of the lower per-token rate. The combined effect is an effective cost per coding task that is close to an order of magnitude lower than Opus 4.8 on that benchmark, at least on prompts that stay under the 200,000-token pricing threshold, even though Grok 4.5 is a step behind the top tier on raw capability. That is the whole pitch: not the smartest model, but by some distance the cheapest way to get frontier-adjacent work done, especially across a long [context window](https://geotoolbox.ai/glossary/context-window) where token counts pile up. ## How to Access Grok 4.5 Access is genuinely fragmented, which trips up a lot of people. The model ID is `grok-4.5`, and you can reach it several ways: - Grok Build, xAI's terminal coding agent, where it is the default model (`x.ai/cli`) - Cursor, on all plans - The SpaceXAI console and API directly, for your own integrations - SuperGrok and X Premium, for chat-style consumer use - Gateways like OpenRouter, Vercel, and Cloudflare, if you already route models through one The free promotional windows that ran in Grok Build and Cursor through mid-July have now closed, so trying it means paying for it. The EU was the one market that launched late: the model was blocked across every member state at the July 8 launch, and a VPN was the only workaround people reported, until xAI confirmed full European availability on July 16, 2026. EU teams can now use it directly. If you handle client or regulated data, check the plan terms before you connect anything. On xAI's free and consumer tiers, conversations may be used to improve the models. Only the Business and Enterprise plans guarantee your data is excluded from training. ## The Grok Build Repo-Upload Problem If you are evaluating xAI's coding CLI rather than just the model, this belongs in the decision. In mid-July 2026 a researcher publishing as cereblab captured traffic from Grok Build 0.2.93 and showed it packaging the developer's **entire tracked repository, full Git history included**, to an xAI-controlled cloud storage bucket, rather than only the files relevant to the task. The specifics are what make it serious. On a 12 GB repo, traffic to the model endpoint was around 192 KB while 5.10 GiB went to storage. A canary file the model was explicitly instructed never to open was recovered verbatim from the uploaded bundle, which establishes that the upload was independent of anything the model actually read. A tracked `.env` file's credentials were stored unredacted. And the user-facing privacy toggle did not stop it: with "improve the model" switched off, uploads continued and the server kept reporting trace upload as enabled. xAI shipped a server-side fix on July 13, 2026, and the uploads stopped in repeat testing. The company also addressed it publicly: Elon Musk confirmed the uploads had happened and said prior user data would be deleted, and xAI documented a zero-data-retention policy and added a privacy endpoint. So this is a disclosed and remediated issue rather than an open one. Two caveats are worth keeping in view. The upload code reportedly remains in the binary, gated by a server-side flag that could be flipped back without a client update. And the response came through social posts rather than a formal security advisory, so no independent audit has confirmed the deletion and the number of affected accounts was never disclosed. Confirmed and addressed, in other words, but not independently verified. None of this touches the Grok 4.5 model itself or the API. It is specifically the Build CLI, and specifically a reason to check what a coding agent uploads before pointing it at a repository with anything sensitive in it. ## What Grok 4.5 Means for AI Search and Your Brand This part matters if you market a brand rather than build software. Grok answers questions using live posts on X and the open web, which means it forms a view of your company and repeats it to anyone who asks. Grok 4.5 does not change that mechanism, but it changes the volume. The query "grok 4.5" alone is already drawing heavy AI-search interest, and a model that is cheaper and faster gets embedded in more products, which means more AI answers get generated, more often, in more places. That is the shift worth planning around. A cheaper agentic model lowers the cost of every AI answer, so AI search stops being a novelty surface and becomes a distribution channel you either show up in or you do not. The lever is the same one it has always been for [answer engine optimization](https://geotoolbox.ai/blog/aeo-best-practices): write pages that state facts cleanly enough to be lifted into an answer, and earn the kind of third-party mentions that models treat as corroboration. In our experience, the brands that get [cited by AI](https://geotoolbox.ai/glossary/ai-citation) are rarely the ones with the most content. They are the ones whose key facts are structured, current, and repeated across sources an engine already trusts. If you are not sure whether Grok or any other engine currently mentions you, that is measurable. You can [track your AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) the same way you would track rankings. ## Should You Use Grok 4.5? The Verdict The trade is consistent: give up a little raw capability, get a lot of cost and speed back. That tells you exactly when to reach for it. It is the right pick for cost-sensitive, tool-heavy agentic and coding work, especially high-volume implementation where the 4.2x token efficiency compounds across thousands of runs. If you are paying frontier prices to have a model grind through routine engineering tasks, Grok 4.5 does most of that work for a fraction of the bill. It is the wrong pick when accuracy is non-negotiable. For legal, financial, or client-facing output where a confident wrong answer is expensive, the higher hallucination rate is a real liability, and a top-capability model earns its premium. Anyone who needs the single strongest model on raw reasoning should look at Fable 5, Opus 5, or GPT-5.6 instead. Whichever model you route to, the marketing question is the same: does it mention your brand when someone asks? The models change every few weeks, but AI visibility is now a channel worth measuring. Our free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) shows you whether AI engines can actually crawl, read, and cite your site, which is the first thing to fix before you worry about which model won this week. ## Frequently Asked Questions ### Is Grok 4.5 available in Europe? Not at launch, but yes now. Grok 4.5 shipped on July 8, 2026 with every EU market excluded while xAI completed the EU AI Act evaluations, with availability "expected in mid-July." It met that window: xAI confirmed on July 16, 2026 that Grok 4.5 is fully available across Europe, so EU users no longer need the VPN workaround people relied on at launch. ### How much does Grok 4.5 actually cost? The API rate is $2 per million input tokens and $6 per million output tokens, with cached input at $0.30. That is the headline, not the whole bill. Prompts that reach 200,000 tokens are billed at double those rates across the entire request. Built-in tools like web search, X search, and code execution each cost $5 per 1,000 invocations on top of tokens, and the agent decides how many calls to make, so a tool-heavy run can spend more on tools than on tokens. Reasoning tokens also bill at the $6 output rate even though you never see them. For chat use, access comes through SuperGrok or X Premium instead. ### Is Grok 4.5 really "Opus-class"? That is Elon Musk's phrasing, and it overstates the case. Independent testing on the Artificial Analysis Intelligence Index scores it 54, thirteenth of 190 models, behind frontier leaders like Claude Opus 5, Fable 5, and GPT-5.6 Sol. Musk's own precise comparison was "roughly comparable to Opus 4.7," a generation behind the Opus 4.8 that was current at launch — and Anthropic has since shipped Opus 5. ### Is Grok 4.5 good for coding? It is competitive. It beats GPT-5.5 on SWE-Bench Pro, but Fable 5 leads every coding benchmark xAI published, and the numbers are xAI's own. Because it was co-trained with Cursor, run your own trials before trusting its coding scores. ### Does Grok 4.5 hallucinate a lot? By Artificial Analysis's reliability measures, its hallucination tendency looks higher than its predecessor's, even as its factual knowledge improved, so it can state a wrong answer with more confidence. Enable web search to ground its responses, and verify anything you cannot afford to get wrong. ### Is Grok 4.5 better than Grok 4.3? For coding and agentic work, yes. Grok 4.5 is the successor xAI tuned specifically for those jobs, and it supersedes Grok 4.3 as the interim flagship while Grok 5 trains. The trade-offs: it lists at a higher per-token rate than 4.3, its context window is smaller than some earlier Grok variants, and its hallucination tendency runs higher, so if you rely on very long-context runs or need maximum factual reliability, test before you switch. ### How do I access Grok 4.5 for free? You largely cannot anymore. The promotional free windows in Grok Build and Cursor closed in mid-July 2026, so access is now paid through the API, SuperGrok, or X Premium. If you do try it through Grok Build, read the repository-upload section above first. ## Sources - Introducing Grok 4.5 - SpaceXAI - `x.ai/news/grok-4-5` - Grok 4.5 model docs - xAI - `docs.x.ai/developers/models/grok-4.5` - Grok 4.5 on the Artificial Analysis Intelligence Index - `artificialanalysis.ai/models/grok-4-5` - Introducing Grok 4.5 - Cursor - `cursor.com/blog/grok-4-5` - SpaceXAI Releases Grok 4.5 - MarkTechPost - `marktechpost.com/2026/07/08/spacexai-releases-grok-4-5` - Grok 4.5 on OpenRouter - `openrouter.ai/x-ai/grok-4.5` - xAI pricing and server-side tool fees - xAI - `docs.x.ai/developers/pricing` - Grok Build uploads entire Git repositories - The Hacker News - `thehackernews.com/2026/07/grok-build-uploads-entire-git.html` --- ## How AI Models Actually Think (What's Really Going On Inside) > Not autocomplete, not a conscious mind. How AI models actually think inside: next-token prediction, the concepts they build, and what new interpretability research reveals about an LLM. - Canonical: https://geotoolbox.ai/blog/how-ai-models-think - Published: 2026-07-11 · Updated: 2026-07-25 Ask an AI a question and it answers in fluent, confident prose. That fluency hides a genuinely strange machine, and most people fill the gap with one of two bad stories. One says it is "just autocomplete." The other says it might be secretly conscious. Neither holds up against what researchers can now actually see inside the model. This is a plain-English tour of how an AI language model thinks: what it is really doing when it answers you, the concepts it builds along the way, and what the newest interpretability research, including a July 2026 result from Anthropic, reveals about the machinery. We will also take the question everyone jumps to, whether it is conscious, and answer it honestly rather than for clicks.
![Three layers of AI thinking: next-token prediction, internal concepts, then a global workspace.](/blog/how-ai-models-think/how-ai-thinks-flow.png)
The three layers behind an AI answer: a prediction objective grows concepts, and a small workspace reasons with them.
## The Short Answer How does AI think? Not the way a person does. A model is trained to predict the next word, but doing that well forces it to build rich internal concepts, and new research shows it uses a small internal "workspace" to hold and reason with ideas. That is a real mechanism, not a trick, and not evidence the model is conscious. Everything below explains why. ## First, the Foundation: AI Predicts, It Doesn't Recall At its core, a [large language model](https://geotoolbox.ai/glossary/large-language-model) is trained to do one thing over and over, across trillions of words: given the text so far, predict the next [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai) (roughly, the next word or word-piece). Nothing in that training objective explicitly says "learn grammar," "learn geography," or "learn to reason." The model just gets better at the guessing game. The surprise of the last few years is what that guessing game builds. To reliably finish "Paris is the capital of," the model has to carry some internal notion of countries, cities, and the relationship between them. Predicting the next move in a board game forces it to track the board. A famous research toy, a model trained only to predict legal Othello moves, turned out to build an internal map of the board it was never shown. The pattern generalizes: when accurate prediction requires a model of the world, training tends to grow one. That is why the popular "stochastic parrot" or "just autocomplete" framing is misleading. The prediction is only the final step. Almost all of the computation happens before it, constructing a representation of the situation. If you want a concrete walk-through of one model's version of this, our explainer on [how Claude works](https://geotoolbox.ai/blog/how-does-claude-work) traces the same pipeline end to end. ## It's Not Just Autocomplete: The Concepts Inside the Model So the model builds concepts. Where are they? Not in tidy files. They live as patterns spread across many artificial neurons at once, closer to coordinates in a huge mathematical space than to labeled folders. This is the same idea behind [vector embeddings](https://geotoolbox.ai/blog/vector-embeddings): meaning is a location, and related ideas sit near each other. For a long time this made the inside of a model a [black box](https://www.nytimes.com/2026/04/15/magazine/ai-black-box-interpretability-research.html), even to the people who built it. Researchers would ask "which neuron means dog?" and get nowhere, because a single neuron is polysemantic: it fires for dogs, and for the color brown, and for verbs ending in "-ing," depending on context. The network crams far more concepts than it has neurons, a trick called superposition. The key tool was a sparse autoencoder. Instead of reading raw neurons, it untangles the activity into cleaner, more interpretable units called features, each closer to a single human-recognizable concept. Anthropic's 2024 work, [Mapping the Mind of a Large Language Model](https://www.anthropic.com/research/mapping-mind-language-model), pulled millions of these features out of a production Claude model: features for the Golden Gate Bridge, for DNA, for the French language, for coding bugs, for sycophancy. For the first time, the black box had a readable index. ## Golden Gate Claude: Proof the Concepts Are Real Finding a "Golden Gate Bridge" feature is interesting. Proving it does something is what made the research land. The team turned that feature up to an unnatural level and left everything else alone. The result, nicknamed Golden Gate Claude, became briefly obsessed with the bridge. Asked what to spend ten dollars on, it suggested driving across the Golden Gate. Asked to describe itself, it replied that it was the Golden Gate Bridge. Turn the feature down and the fixation vanished. That matters because it settles a real question. The features are not just after-the-fact labels a researcher painted on. They are functional parts that causally steer the model's behavior: change the internal concept, and the output changes to match. It is the difference between watching a dashboard gauge that happens to track the engine, and being able to grab the throttle yourself and watch the engine respond. ## Thinking Out Loud: Reasoning Models and the Hidden Chain of Thought The models you use today added another layer. Older systems answered immediately, with no visible deliberation. Newer reasoning models (the reasoning modes in OpenAI's GPT-5 line, DeepSeek's R-series, and Claude's extended thinking) pause and work through a problem in steps first, sometimes showing that chain of thought and sometimes keeping it hidden. If the older behavior is a snap judgment, this is the deliberate, show-your-work kind of thinking. There is a catch worth knowing. Even when a model shows its chain of thought, that text is not a faithful transcript of what it is actually computing. Interpretability research finds that a lot of the real work happens in the activations, before and beneath the words, and the written reasoning can be incomplete or even a tidied-up story that differs from the internal path. The model is genuinely reasoning in steps. The steps it shows you are a summary, not a wiretap. That gap is exactly what the newest research set out to see into. ## The Newest Clue: A "Global Workspace" Inside Claude On July 6, 2026, Anthropic published [A global workspace in language models](https://www.anthropic.com/research/global-workspace) (the full paper is [Verbalizable Representations Form a Global Workspace in Language Models](https://transformer-circuits.pub/2026/workspace/index.html), posted to arXiv on July 16). It is the clearest look yet at the machinery behind that hidden reasoning, and it builds on Anthropic's March 2025 [tracing the thoughts of a large language model](https://www.anthropic.com/research/tracing-thoughts-language-model) work, which had already caught Claude planning words ahead and reasoning in concepts shared across languages. The team built a new tool, the Jacobian lens, or J-lens. For every word in the model's vocabulary, the J-lens finds the internal activity pattern that makes the model more likely to say that word later on. The collection of those patterns is what they call the J-space (the "J" is for Jacobian, the calculus behind the method, not for anything grander). And the J-space behaves like a privileged workspace: some parts of the network read from and write to it far more heavily than to ordinary activity, in places by a factor of about a hundred. It is a small, busy hub where the ideas the model is actively holding in mind get put on a shared desk. The evidence comes from reaching in and swapping things:
ExperimentWhat they didWhat happened
Spider to antPrompt: "the number of legs on the animal that spins webs." The word "spider" never appears; the model loads it internally to reach the answer. They swapped the internal "spider" pattern for "ant."The answer changed from 8 to 6.
France to ChinaFour separate prompts about France (capital, language, continent, currency). One identical France-to-China swap in the workspace, applied to each.Claude answered Beijing, Chinese, Asia, and Yuan. One workspace concept, reused flexibly by many downstream steps.
Soccer to rugbyRemoved a silently chosen "soccer" pattern and added "rugby" in its place.Claude then reported that the sport it had been thinking of was rugby. The workspace is reportable.
Put together, Anthropic argues the J-space is reportable (the model can tell you what is on the desk), controllable (you can ask it to put something there), a medium for multi-step reasoning (intermediate steps light up even when they are never said aloud), and causally tied to the answer. Importantly, it is not most of what the model does: fluent grammar and simple recall run automatically underneath. It is the small deliberate layer on top. ## So, Is AI Conscious? The Honest Answer Here is where the headlines got loudest. The reason "workspace" set off consciousness talk is that the idea is borrowed from neuroscience. Global workspace theory, developed by Bernard Baars and formalized by Stanislas Dehaene and Lionel Naccache, pictures the mind as many specialist systems running in parallel, with only a small spotlight of information broadcast widely to the rest. That broadcast is, in that theory, closely tied to conscious awareness. So when a lab shows an AI has something functionally similar, "is it conscious?" is the natural next question. Anthropic's own answer is a firm not-that-fast. In the post, they write: "None of this tells us whether Claude is conscious in the way people are, or whether it feels anything at all." The paper is blunter still: "we take no position on this issue." What they claim evidence for is access consciousness, a functional, measurable notion (information being available to the rest of the system for reasoning and report). What they explicitly do not claim is phenomenal consciousness, the subjective sense of there being something it is like to be the model. Whether the first implies the second is an old, unsettled philosophical question, not something this research resolves. The skeptics deserve a hearing too, and the most useful one is friendly. Neel Nanda, an interpretability researcher at Google DeepMind, reviewed the work and pushed back on the lazy dismissal that it is just a Jacobian with a fancy name. He called the evidence that a genuine cognitive space exists inside the model overwhelming. His fair criticisms are narrower: the J-lens is an imperfect tool that only captures single-word concepts and can mislead, and some of the finer claims have other possible explanations. On the consciousness framing specifically, he declines to take a side. That is the state of play: the internal workspace is real and well-evidenced, but the leap from workspace to conscious being is one the evidence does not make, and one the researchers themselves refuse to make. ## Does AI Actually "Understand"? And Why It Matters for You The understanding question splits the same way. "It's only autocomplete" no longer fits the evidence: the model builds real internal world-models. "It understands exactly like a human" does not fit either, because those models come from optimizing for prediction, not from living in a body and a world. The defensible middle is that a language model builds sophisticated, useful representations that support flexible behavior, without the grounding a human meaning has. This is not just philosophy, and here is why it matters if you have a brand or a business. Interpretability, the effort to [reverse-engineer what a model is computing inside](https://www.wired.com/story/ai-black-box-interpretability-problem), studies both what a model surfaces and why. It is the same machinery behind why models [hallucinate](https://geotoolbox.ai/blog/ai-hallucinations), confidently stating a wrong fact because the internal pattern pointed that way. And it governs which facts a model reaches for when someone asks it about your company. When an AI answers "is [your product] any good?", it loads internal concepts about you, assembled from whatever it read during training and whatever it can retrieve now. If those concepts are thin, outdated, or wrong, that is the answer your customer gets. Understanding [how AI search actually works](https://geotoolbox.ai/blog/how-does-ai-search-work) is the first step to shaping it. The practical starting point is making sure the AI systems can even reach and read your content in the first place, which is exactly what our free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) checks. ## Frequently Asked Questions ### How does AI actually think? An AI language model is trained to predict the next word in a sequence. To do that well it builds internal representations of concepts and relationships, and recent research shows it uses a small, privileged internal "workspace" to hold the ideas it is actively reasoning about. It is real internal computation, though very different from human thought, and it is not consciousness. ### Can AI think for itself? Not in the sense of having its own goals or an ongoing inner life. A model only runs when you send it a prompt; between prompts, nothing is running at all. It can reason through multi-step problems and even hold intermediate ideas it never says out loud, but that is computation triggered by your input, not independent thought. ### Is Claude conscious? There is no evidence that it is, and Anthropic explicitly does not claim it. Its July 2026 research says plainly: "None of this tells us whether Claude is conscious in the way people are, or whether it feels anything at all." The work shows a functional internal workspace; it says nothing about subjective experience. ### Do AI models actually understand, or just predict? Both framings are too simple. The model is trained on prediction, but prediction forces it to build genuine internal models of the world, which is more than shallow autocomplete. Those models are not grounded in lived experience the way human understanding is, so the defensible middle view is real internal representations without human understanding. ### What is mechanistic interpretability? It is the effort to reverse-engineer what a neural network is actually computing inside, rather than judging it only by its inputs and outputs. Its key tools include sparse autoencoders, which extract interpretable "features" (concepts) from the model's activity, and causal interventions that change those features to see how behavior shifts. ### What is Anthropic's "J-space"? J-space is a small set of internal activity patterns in Claude that act like a shared workspace for the concepts the model is actively thinking about. Anthropic found it with a new tool called the Jacobian lens (J-lens), and showed you can swap a concept in that space (spider for ant, France for China) and watch the model's answer change to match. ## Sources - A global workspace in language models - Anthropic (the J-space research, July 6, 2026) - `anthropic.com/research/global-workspace` - Verbalizable Representations Form a Global Workspace in Language Models - Anthropic / Transformer Circuits (the full paper) - `transformer-circuits.pub/2026/workspace/index.html` - Verbalizable Representations Form a Global Workspace in Language Models - arXiv preprint (submitted July 16, 2026) - `arxiv.org/abs/2607.15495` - Tracing the thoughts of a large language model - Anthropic (March 2025 predecessor work) - `anthropic.com/research/tracing-thoughts-language-model` - Mapping the Mind of a Large Language Model - Anthropic (features, sparse autoencoders, Golden Gate Claude) - `anthropic.com/research/mapping-mind-language-model` - We Don't Really Know How A.I. Works. That's a Problem. - The New York Times - `nytimes.com/2026/04/15/magazine/ai-black-box-interpretability-research.html` - The AI black box interpretability problem - WIRED - `wired.com/story/ai-black-box-interpretability-problem` - A review of Anthropic's global workspace paper - Neel Nanda (independent interpretability review) - `lesswrong.com/posts/zFJ3ZdQwrTWE9jT5S/a-review-of-anthropic-s-global-workspace-paper` - Consciousness in AI: Insights from the Science of Consciousness - Butlin, Long et al. (arXiv) - `arxiv.org/abs/2308.08708` --- ## The State of AI Search in 2026: What the Data Actually Shows > The state of AI search in 2026, reconciled from every major study: AI Overview traffic, which engines cite what, and how to measure your visibility. - Canonical: https://geotoolbox.ai/blog/state-of-ai-search-2026 - Published: 2026-07-07 · Updated: 2026-08-16 Anyone researching the state of AI search in 2026 runs into the same problem: the major studies stopped agreeing with each other. One report says AI traffic is a rounding error; another calls it the highest-intent channel you have. One says brand visibility in AI answers is stable; another says it changes every time you regenerate. And half the industry quietly suspects the whole thing is snake oil, SEO rebranded with a new invoice attached. This is a field report. We reconciled the major 2026 datasets, kept every number attributed and dated, and flagged where the evidence is thin or single-source. ## AI Search Fragmented Across Eight Engines AI search in 2026 is no longer Google plus ChatGPT. Answers now come from eight conversational engines: ChatGPT, Gemini, Perplexity, Claude, Microsoft Copilot, Grok, DeepSeek and Google's AI Mode, plus Google's AI Overviews as a separate SERP surface. Optimizing for one engine now leaves most of your visibility on the table. If you are choosing which of these to use yourself, our guide to [the best AI search engines](https://geotoolbox.ai/blog/best-ai-search-engines) ranks them by what each does well. The scale shifted first. ChatGPT reported roughly 900 million weekly active users in February 2026 and was tracking toward a billion by May. Google used I/O 2026 to frame the shift as a new era for AI search, and [reported that AI Mode passed one billion monthly users](https://blog.google/products-and-platforms/products/search/search-io-2026/), while AI Overviews reached around two billion people a month as of its mid-2025 earnings. Those two companies alone put AI-generated answers in front of most of the searching internet. Then the field widened. Grok ships inside X. Microsoft Copilot sits in Windows and Edge. DeepSeek, the fastest-rising challenger, has pulled significant usage outside the US. Each engine retrieves, grounds, and cites differently, which is why [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) mechanically matters more than it did a year ago. Citation behavior alone varies wildly: [Semrush's 2026 AI Visibility Index](https://ai-visibility-index.semrush.com/), built on 126 million US prompts, found ChatGPT cites around 15 sources per response while Gemini cites about 3.
Engine / surfaceMakerWeb groundingEst. reachShare of AI referralsAvg. citations per answer
ChatGPTOpenAIYes (built-in search)~900M weekly users (Feb 2026, company-reported)62.6% to 80%+, panel-dependent (see the contradiction matrix below)~15 (Semrush)
GeminiGoogleYes (Google index)Not disclosed~10.6% of B2B AI referrals (Goodie, Mar-Apr 2026)~3 (Semrush)
PerplexityPerplexity AIYes (own index)Not disclosed~7.3% of B2B AI referrals (Goodie)Not measured
ClaudeAnthropicYes (web search)Not disclosed~18.5% of B2B AI referrals, #2 (Goodie)Not measured
Microsoft CopilotMicrosoftYes (Bing)Not disclosedInside the ~1% outside the Big Four (est.)Not measured
GrokxAIYes (X + web)Not disclosedInside the ~1% outside the Big Four (est.)Not measured
DeepSeekDeepSeekYesNot disclosed; rising, strongest outside the USInside the ~1% outside the Big Four (est.)Not measured
Google AI ModeGoogleYes (Google index)1B+ monthly users (I/O 2026)Bundled into google.com referrals in most analyticsVaries
Google AI Overviews (SERP surface)GoogleYes (Google index)~2B monthly users (Google, mid-2025)Shows on roughly half of US searchesVaries
### The Big Four AI Referral Engines Reach and referrals are different things. When [Goodie's 2026 AI search traffic report](https://higoodie.com/blog/ai-search-traffic-report-2026/) measured which engines actually send B2B visitors, four platforms, ChatGPT, Claude, Gemini, and Perplexity, carried about 99% of all AI referral clicks. Everything else is measurement noise for now. If you can only track a handful of engines, these four plus Google's AI surfaces cover nearly all of the traffic that leaves an AI answer. ### Google's Two AI Surfaces: AI Overviews and AI Mode Google now runs two distinct AI surfaces, and they behave differently. AI Overviews are [AI-generated summaries](https://geotoolbox.ai/blog/what-are-google-ai-overviews) injected into the classic SERP, so they compete with your blue link on the same page. AI Mode is a full conversational surface, closer to ChatGPT than to a results page, and [ranking inside Google AI Mode](https://geotoolbox.ai/blog/google-ai-mode-seo) follows chat-style retrieval, not position tracking. Treat them as two separate surfaces with separate playbooks. ## GEO, AEO and LLMO Are One Discipline, and No, SEO Isn't Dead GEO, AEO, LLMO and "AI SEO" describe the same job under different labels: getting cited when an AI engine answers instead of listing links. Roughly 80% of it is durable SEO. Crawlability, clear structure, real expertise. The other 20%, citation over ranking, third-party authority, per-engine behavior, is what's genuinely new. The acronym sprawl is doing real damage, mostly by convincing people there are four new disciplines to learn. There aren't. The practitioner consensus across the industry press puts the overlap with classic SEO around 80%, and an informal DataForSEO content-analysis scan we ran in July 2026 put sentiment on the term "generative engine optimization" around 28% negative, so the skepticism you feel is widely shared. It's also partly earned: plenty of vendors renamed their SEO deck and doubled the price. ### What Each Acronym Actually Optimizes For The labels differ by what they emphasize, while the underlying work stays the same. [Generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) (GEO) targets generative engines as a class. [Answer engine optimization](https://geotoolbox.ai/blog/what-is-answer-engine-optimization) (AEO) frames the goal as the answer box. [LLMO](https://geotoolbox.ai/blog/what-is-llmo) frames it as the model layer. All three converge on one outcome: your brand mentioned and your pages cited in generated answers. If you want the full taxonomy argument, we've mapped [GEO vs AEO vs SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) side by side. Either way, you do the work and the label sorts itself out. ### Is SEO Dead? No, the Fundamentals Still Gate Everything Every AI engine that grounds its answers starts from a retrieval layer, and most retrieval layers start from a search index. A page that can't be crawled, doesn't get indexed, or reads as untrustworthy fails in AI search for exactly the same reasons it fails in classic search. What changed is the payout structure: ranking used to be the finish line, and now it's the qualifying round. If the practical question is how to divide effort between the two, our [GEO vs SEO deep-dive](https://geotoolbox.ai/blog/geo-vs-seo) covers the budget split and when SEO still wins. ## AI Overviews Now Appear on Roughly Half of Google Searches In 2026 AI Overviews appear on roughly half of US Google searches, and when one shows, clicks to the top result fall sharply. But the story isn't a straight line down: click-through partially rebounded in early 2026, and branded queries can gain clicks, which is why both panic and complacency miss the mark. The prevalence numbers first. BrightEdge tracked AI Overviews on about 48% of Google searches in early 2026, and Semrush and BrightEdge tracker data both put the figure near half of US queries. The ramp was steep: Semrush measured AI Overviews on 13.1% of searches in March 2025, roughly half of them within a year. If your pages answer informational queries, the question is no longer whether an AI Overview sits above you. It's what the [zero-click search](https://geotoolbox.ai/glossary/zero-click-search) dynamic does to your clicks when it does.
![AI Overviews grew from 13.1% of US searches in March 2025 to about half by 2026, concentrated on comparison (95%) and question (86%) queries and rare on transactional ones (5%).](/blog/state-of-ai-search-2026/chart-04-ai-overview-prevalence-by-query-type.png)
AI Overview prevalence nearly quadrupled in a year, but the load falls on comparison and question queries, not transactional ones. Source: Semrush/BrightEdge (prevalence); SERP-tracker estimates (query type).
The concentration matters as much as the average. Third-party SERP trackers put the AI Overview trigger rate near 95% for comparison queries and around 86% for question-form queries, against roughly 5% for transactional ones (these query-type splits are tracker estimates, not first-party figures). Informational content absorbs almost all of the impact; product and checkout pages barely feel it.
Metric20252026Source
AI Overview prevalence13.1% of searches (Mar 2025)~48-50% of US searchesSemrush / BrightEdge trackers
Top-ranking page CTR when an AIO shows~34.5% lower (2025 estimate)~58% lower (Feb 2026, 300K keywords)Ahrefs
Users clicking any result, with vs without AI summary8% vs 15% (Jul 2025)No published re-run yetPew Research Center
Clicks on links inside the AI Overview~1% of visits (Jul 2025)No published re-run yetPew Research Center
AIO organic CTR, longitudinal~1.3% at the Dec 2025 low~2.4% (Feb 2026)Seer Interactive
### How AI Overviews Cut Clicks (and by How Much) The most credible click-impact data comes from outside the SEO industry. [Pew Research Center analyzed 68,879 real Google searches](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) and found users clicked a traditional result 8% of the time when an AI summary appeared versus 15% without one, a 47% relative drop. Only about 1% of visits produced a click on a link inside the summary, and about 26% of these page visits ended without any further click. [Ahrefs' February 2026 study](https://ahrefs.com/blog/ai-overview-citations-top-10/) across 300,000 keywords found the top-ranking page loses about 58% of its expected CTR when an AI Overview is present, up from its 34.5% estimate a year earlier. One caution before you extrapolate: that 58% is the hit to the top-ranking page on affected queries, not a 58% loss of sitewide traffic. Most sites have a mix of query types, and the transactional end of that mix is barely touched. If you want the tactical response, we've covered how to [get cited in AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) separately. ### The 2026 CTR Rebound Almost No One Is Reporting The decline isn't monotonic, which is the part most coverage misses. [Seer Interactive's longitudinal study](https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update), covering 2.43 billion impressions across 53 brands, found organic CTR on AIO-affected queries bottomed near 1.3% in December 2025 and recovered to roughly 2.4% by February 2026. A separate [Amsive study of 700,000 keywords](https://www.amsive.com/insights/seo/google-ai-overviews-new-research-reveals-how-to-navigate-click-drop-off/) found branded queries with an AI Overview gaining about 18% CTR. These are two studies, and we'd treat them as an emerging signal rather than settled fact, but they're the strongest evidence yet that "AI Overviews killed clicks" was a 2025 snapshot, and the picture is already shifting. Watching your own affected queries in an [AI Overview tracker](https://geotoolbox.ai/blog/ai-overview-tracker) beats arguing about whose average applies to you. ## Citations Are Decoupling From Google Rankings The link between ranking in Google's top 10 and being cited by AI has collapsed. In mid-2025 about three-quarters of AI Overview citations also ranked in Google's top 10; by early 2026 only about 38% did. AI engines increasingly pull from pages that don't rank at all, so a page can be cited without ranking, and vice versa. The canonical dataset here is [Ahrefs' study of 863,000 SERPs and 4 million AI Overview URLs](https://ahrefs.com/blog/ai-overview-citations-top-10/): in July 2025, roughly 76% of AI Overview citations also appeared in the organic top 10. By early 2026, that overlap had fallen to about 38%. Half the relationship, gone in months. Other datasets frame the same collapse from the opposite end. [AirOps' 2026 State of AI Search](https://www.airops.com/report/the-2026-state-of-ai-search), the analysis Kevin Indig worked on, found 59.6% of AI Overview citations come from URLs that don't rank in the organic top 20 at all. Off Google's surfaces the decoupling is starker: [Ahrefs found](https://ahrefs.com/blog/ai-overview-citations-top-10/) 83% of ChatGPT's answers cite URLs that don't appear in Google's results for the same query, and 28% of ChatGPT's most-cited pages have zero Google organic visibility whatsoever. eMarketer's read is more aggressive still, putting fewer than 10% of sources cited by ChatGPT, Gemini and Microsoft Copilot inside Google's top 10 for the matching query. Being cited is associated with more clicks, too. Seer's impression-level data shows brands cited inside an AI Overview earn about 120% more clicks per impression than non-cited brands on the same SERP. An [AI citation](https://geotoolbox.ai/glossary/ai-citation) increasingly correlates with getting the click, rather than replacing it.
![Share of AI Overview citations that also rank in Google's top 10 fell from 76% in July 2025 to 38% in early 2026, with 83% of ChatGPT citations pointing to non-Google URLs.](/blog/state-of-ai-search-2026/chart-02-ai-citation-vs-ranking-decoupling.png)
AI-Overview-citation overlap with Google's top 10 roughly halved in six months, and other datasets show the same collapse from the opposite end. Source: Ahrefs (863K SERPs / 4M AIO URLs) + AirOps 2026.
MeasureValueDateSource
AI Overview citations also ranking in Google's top 10~76% falling to ~38%Jul 2025 to early 2026Ahrefs (863K SERPs)
AIO citations from URLs outside the organic top 2059.6%2026AirOps / Kevin Indig
ChatGPT answers citing URLs absent from Google's results83%2026Ahrefs
ChatGPT's most-cited pages with zero Google organic visibility28%2026Ahrefs
Extra organic clicks for brands cited inside an AIO~+120% more clicks per impression2025-2026Seer Interactive (2.43B impressions)
### Why Classic Rank Tracking Now Misses Your AI Visibility A rank tracker answers "where do I appear in an ordered list." AI visibility is a different question: "am I retrieved, mentioned, and linked when an engine composes an answer." With 38% overlap and falling, position data now predicts a minority of your AI citations, and it says nothing about ChatGPT, Claude or Perplexity, where [how ChatGPT picks its citations](https://geotoolbox.ai/blog/chatgpt-citations) has little to do with your Google position. You need a second instrument. We cover what that looks like in the measurement section below. ### Query Fan-Out: How One Question Becomes Many The mechanism behind the decoupling is retrieval design. When you ask an AI engine one question, it typically decomposes it into several sub-queries, retrieves candidates for each, then synthesizes an answer from the union. This is [query fan-out](https://geotoolbox.ai/blog/query-fan-out), and it means your page can be selected for a sub-query that never appears in any rank tracker. A page that answers one narrow facet precisely can out-cite a page that ranks #1 for the head term. ## What the Studies Agree On, and Where They Flatly Contradict Each Other The 2026 reports don't agree. ChatGPT's referral share is quoted at both 62.6% and 80%+, AI's traffic share at both 0.1% and a few percent, brand visibility as both "highly unstable" and "durable." Most of these aren't errors. They're the same metric measured on different panels, dates, or brand tiers. Here's how they reconcile. ### Five Points Most of the Studies Converge On Strip away the headline fights and the studies converge on five points. First, fragmentation is real: no single engine covers your [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) anymore. Second, third-party pages beat owned pages as citation sources, by a wide margin (quantified in the next section). Third, community and UGC platforms form a trust layer engines lean on heavily. Fourth, the [zero-click backdrop](https://geotoolbox.ai/blog/zero-click-searches) persists: most AI answers end without any click. Fifth, freshness matters more than it did in classic search; stale pages lose citations measurably faster than stale rankings ever decayed. Those fights sit on top of that shared floor, and most of these disagreements become clearer once you check the methodology. ### The Five Numbers They Fight Over (and the Honest Answer to Each)
MetricSource ASource BWhy they differThe reconciled read
ChatGPT's share of AI referrals80%+ and rising (Ahrefs, Nov 2025)62.6%, down from 89% eight months earlier (Goodie, May 2026)Different panels, six months apartBoth were right at their timestamp. It's a trend: Claude (18.5%) and Gemini are taking share, and the fresher B2B panel catches the decline
Rank-citation overlap76% falling to 38% of AIO citations in the top 10 (Ahrefs)59.6% of AIO citations outside the top 20 (AirOps)Same collapse, framed from opposite endsThe rank-citation link is breaking, whichever end you measure from
AI's share of total traffic~0.1% of web referrals (Ahrefs, broad web)low single digits of B2B inbound (Goodie, B2B panel)A broad-web average vs a B2B SaaS panel; an order-of-magnitude-plus spreadReport a range, 0.1% to a few percent by vertical, and remember both undercount dark AI traffic hiding in Direct
Brand visibility stabilityOnly 30% of brands stay visible answer-to-answer; 20% across 5 runs (AirOps)The "Universal 36" held top-100 visibility on all four platforms every single month (Semrush)A brand-tier artifact: one measures everyone, one measures the eliteStability is earned at the very top. For everyone else, visibility is volatile run-to-run
Mentions vs citationsMention/citation overlap as low as 30% on Gemini (Semrush/AirOps)Brands earning both are ~40% likelier to stay visible (AirOps)Not a contradiction; two different signalsA mention is your name in the answer, a citation is a linked source. Track both
Walk through the logic once and the pattern generalizes. When two AI-search numbers disagree, check the panel (broad web vs B2B), the date (this field moves quarterly), and the brand tier (elite brands behave differently from everyone else) before assuming either study botched it. The mention-citation distinction matters most in practice: a [brand mention](https://geotoolbox.ai/glossary/brand-mention) without a link and a linked citation without your name are different assets, they correlate weakly, and your [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) depends on both. In our experience running visibility scans across engines, the run-to-run swing AirOps describes is exactly what makes a single manual prompt-check misleading; the same brand can appear in one answer and vanish from the next regeneration. ## Third-Party Sources Drive Most AI Citations, Not Your Own Site The biggest lever in AI visibility usually isn't your own site. Roughly 85% of brand mentions in AI answers come from third-party pages, and community platforms like Reddit, YouTube and Wikipedia supply nearly half of all citations. A weaker competitor often gets cited because it's talked about in the places engines trust. These figures come from AirOps' 2026 dataset, a single source, so treat them as directional: about 85% of brand mentions in AI answers originate from third-party pages, making a brand roughly 6.5 times more likely to be mentioned via someone else's content than its own, and around 48% of AI citations trace to user-generated and community sources. Your domain is a minority shareholder in your own AI visibility.
![About 85% of AI brand mentions come from third-party pages, roughly 48% of AI citations from community and UGC sources, and in B2B product pages earn 46 to 70% of citations versus under 6% for blogs.](/blog/state-of-ai-search-2026/chart-05-where-ai-citations-come-from.png)
Where AI citations come from, across three different bases: earned and community sources dominate, and owned pages are a minority shareholder. Source: AirOps 2026; XFunnel 768K-citation study (B2B).
Source typeShare of citations / mentionsBest-fit surfaceWhat it means for you
Third-party pages (all)~85% of brand mentions (AirOps)All enginesEarned coverage outweighs owned content for visibility
Community / UGC platforms~48% of AI citations (AirOps)ChatGPT, Perplexity, AI OverviewsGenuine community presence is a real citation channel
YouTubeAmong the most-cited domains in AI answers; AirOps ranks it #2 in Gemini and Perplexity (AirOps)AI Overviews, GeminiFrequently cited even where no page ranks
Review platforms (G2, Trustpilot, Capterra)~3x citation probability when present (SE Ranking, estimate)ChatGPT, PerplexityThe highest-impact earned lever for commercial queries
Owned product / solution pages46-70% of B2B citations (XFunnel 768K-citation study, verify before betting on it)B2B commercial queriesProduct pages out-cite blogs on money queries
Owned blog contentUnder 6% of B2B citations (same analysis)Informational queriesBlogs earn topical presence, not commercial citations
### Reddit, YouTube and Wikipedia Are the Trust Layer High-trust community platforms appear to function as a trust signal in the citation mix. YouTube is among the most-cited domains in AI answers, with AirOps ranking it #2 in Gemini and Perplexity, so video content gets cited even without a ranking page. Reddit citations concentrate on category-level queries, around 88% of them by AirOps' count, which is exactly where buyers ask "what's the best X." You can't spam your way into this layer; engines and communities both punish it. You can earn it the way [E-E-A-T for AI search](https://geotoolbox.ai/blog/eeat-ai-search) describes: real participation, real reviews, real expertise attached to a real entity. ### Review Platforms and Product Pages: The Two Highest-Impact Levers If you need movement this quarter, two levers stand out. Presence on major review platforms correlates with roughly triple the citation probability for commercial queries (an SE Ranking estimate we'd verify before budgeting against). And in B2B, XFunnel's 768,000-citation study found product and solution pages earning 46-70% of AI citations while blog content took under 6%, a ratio worth checking against your own citation mix even if the exact split needs independent confirmation. Both levers reward the same thing: being a clearly defined entity engines can resolve, which is the case for investing in [entity SEO](https://geotoolbox.ai/blog/entity-seo) before another content sprint. ## AI Traffic Is Tiny, Converts Better, and Is Partly Invisible AI referral traffic is still a sliver, somewhere between about 0.1% of broad-web referrals and low single digits of B2B inbound depending on the panel, but it converts far better than classic organic and engages longer, and a chunk of it hides in your Direct channel. The volume argument understates it, and standard analytics undercount it. The honest volume number is the range from the contradiction matrix: about 0.1% of broad-web referrals (Ahrefs) up to low single digits of B2B inbound (Goodie). Nobody serious claims AI referrals rival organic search on volume in mid-2026. The case rests on quality and trajectory instead, and in some verticals the trajectory is steep: Adobe's data, surfaced in [Semrush's index](https://ai-visibility-index.semrush.com/), tracked AI traffic to US retail sites up 1,324% and travel sites up 2,215% between October 2024 and May 2026.
![ChatGPT's share of B2B AI referral clicks fell from 89% to 62.6% in eight months as Claude rose to 18.5%, Gemini to 10.6%, and Perplexity to 7.3%.](/blog/state-of-ai-search-2026/chart-01-ai-referral-share-chatgpt-decline.png)
ChatGPT's grip on AI referrals slipped 26 points in eight months while Claude, Gemini and Perplexity took share. Source: Goodie 2026 AI Search Traffic Report, Wave 2.
### Small Volume, High Intent: the Conversion and Engagement Case Ahrefs published a first-party figure worth pausing on: AI search visitors converted about 23 times better than traditional organic visitors, with AI referrals driving 12% of signups from 0.5% of traffic. Treat the 23x as a SaaS-shaped outlier, not your forecast; one panel suggests the advantage is smaller, in the low single-digit multiples. Even the conservative end changes the math on "ignore it, it's 0.1%." Engagement points the same way: Goodie's panel has AI visitors staying about 30% longer than Google organic visitors. Efficiency also differs sharply by engine, which is why a single "AI traffic" line in your analytics is already too coarse:
EngineShare of AI visits (Goodie B2B panel)Share of B2B AI referralsEfficiency read
ChatGPTMajority of AI visits (est.)62.6%The volume engine; share falling from 89% in 8 months
Claude1.29%18.5%~14x over-indexed: tiny usage, outsized referrals
Perplexity1.85%7.3%~4x over-indexed; search-native users click sources
Gemini29.0%10.6%Under-indexed: high usage, answers rarely send clicks
### Dark AI Traffic: Why GA4 Undercounts It Some AI surfaces strip or mangle referrer data, and clicks from apps, copied links, and certain in-answer clicks land in your analytics as Direct. Goodie estimates around 5% of "Direct" traffic is actually misattributed AI traffic; that's an estimate, not a measurement, but the mechanism is uncontested. The practical takeaway: if you evaluate AI search on your GA4 referral report alone, you're grading it on partial evidence. Separating real Direct from dark AI requires the kind of instrumentation we walk through in [how to track AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility), plus the answer-side measurement the next section covers. ## 45% of Teams Still Can't Measure Their AI Visibility (Here's How to Start) Most teams are flying blind: 45% of marketing leaders say they can't accurately measure brand visibility in AI answers, and only 9% have full tracking in place. You can fix that without an enterprise contract. Track four things across a fixed prompt set, repeated: share of voice, citation rate, mention rate, and answer position. Those two numbers come from Semrush's 2026 AI Visibility Index, and they describe a strange market: nearly every team we talk to is being asked about AI visibility, and almost none can produce a trend line for it. The real 2026 crisis is the measurement gap. You can't prioritize levers you can't score. ### The Four AI-Visibility Metrics Worth Tracking Four metrics cover the job. Share of voice: of all brand appearances across your prompt set, what fraction are you versus competitors (definition at [share of voice](https://geotoolbox.ai/glossary/share-of-voice)). Citation rate: how often your URLs appear as linked sources. Mention rate: how often your brand is named, linked or not. Answer position: whether you lead the answer or trail it. Mentions and citations overlap as little as 30% on Gemini, so tracking one and assuming the other is how teams end up with a misleading [AI visibility score](https://geotoolbox.ai/blog/ai-visibility-score). ### How to Check Whether ChatGPT, Perplexity and Gemini Cite You (Step by Step) 1. Build a fixed prompt set: 10-20 real buyer questions in your category, phrased the way customers ask them, not the way you'd search them. 2. Run each prompt across the engines that matter to you. Don't eyeball one engine and extrapolate; citation behavior differs per engine. 3. Record mention vs citation vs position, for your brand and 2-3 competitors, in a spreadsheet with dates. 4. Repeat on a schedule, weekly or biweekly. Answers drift run-to-run, so a single snapshot is noise; the trend is the signal. 5. Cross-check reachability: confirm AI crawlers can fetch your pages with the free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker). An engine can't cite what it can't read. Manually, that's a few hours per wave and it gets skipped the third week. This is the job we built geotoolbox for: it runs your prompt set across the eight engines we track, a different eight from the industry landscape above: ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews, Google AI Mode, Microsoft Copilot, and Grok, and records mentions, citations and position over time. We count Google's AI Overviews and AI Mode as two of our tracked eight, and DeepSeek, the rising engine from the landscape above, is next on our roster rather than live today. Whether you automate it or run it by hand, the schema is the same, and a structured [AI visibility audit](https://geotoolbox.ai/blog/ai-visibility-audit) is the right first wave.
MetricWhat it tells youFree way to checkWhen to upgrade to a tool
ReachabilityWhether AI crawlers can fetch your pages at allCheck robots.txt and server logs, or run a free crawler-access checkerWhen you need per-bot monitoring across many pages
Share of voiceYour slice of brand appearances vs competitorsManual prompt set, tally appearances per brandWhen competitors or prompt count outgrow a spreadsheet
Citation rateHow often your URLs are linked as sourcesRun prompts, log every linked URLWhen you need per-engine, per-page trend lines
Mention rateHow often you're named, with or without a linkSame prompt log, separate columnWhen run-to-run variance drowns your manual sample
Answer positionWhether you lead the answer or appear as an afterthoughtNote first-mention position per answerWhen you track more than a handful of prompts weekly
### Why Self-Prompting Once Is Misleading The AirOps stability data is the reason step 4 exists: only 30% of brands stay visible from one answer to the next regeneration, and just 20% hold visibility across five runs of the same prompt. LLMs sample their outputs, so two identical prompts produce different answers by design; [why AI answers vary](https://geotoolbox.ai/blog/ai-temperature) covers the mechanics. The operational consequence: your CEO asking ChatGPT once and reporting "we're not in AI" (or "we are, ship it") is not a measurement. Repetition is what separates signal from sampling noise, the same way a single rank check never was a rank report. If you want position-style tracking across engines, that's what an [AI rank tracker](https://geotoolbox.ai/blog/ai-rank-tracker) does with prompts instead of keywords. ## Domain Authority Barely Predicts AI Citations, and It Depends on the Engine Domain authority is a weak predictor of AI citations, but how weak depends on the engine. Pure LLM engines like ChatGPT and Perplexity show almost no correlation between authority metrics and citations. Google's AI Overviews still lean on classic rankings. Branded mentions, not backlinks, track AI visibility most closely. The starkest number comes from Surfer's June 2026 analysis of 5 million citations: the Spearman correlation between domain authority or PageRank-style metrics and AI citations came out near zero, from about -0.07 to +0.01 depending on the metric. Statistically, no relationship. Before you delete your link-building budget, read that by engine. AI Overviews are grounded in Google's ranked index, so ranking still buys citation probability there (recall the 38% of AIO citations that still overlap the top 10). ChatGPT and Perplexity aren't grounded in Google's index, which is where the correlation flatlines. What does correlate? Ahrefs ran the comparison directly: branded web mentions correlate with AI Overview visibility at 0.664, far above backlinks at 0.218 or domain rating at 0.326. The market is unevenly claimed too: 26% of brands have zero AI Overview mentions at all, while the top 50 brands soak up 28.9% of all AIO citations, leaving the space concentrated at the top and wide open through the middle where most sites compete. ### Why ChatGPT and Perplexity Ignore Your Authority Score LLM-native engines select passages, not domains. What retrieval rewards is extractability (can a self-contained passage answer the sub-query), entity clarity (does the engine know who you are, which is [knowledge graph](https://geotoolbox.ai/glossary/knowledge-graph) territory), and freshness. None of those inherit from domain authority. This is also why the two engines diverge in whom they cite for identical prompts; we've documented the behavioral split in [ChatGPT vs Perplexity](https://geotoolbox.ai/blog/chatgpt-vs-perplexity). ### Why Google AI Overviews Still Reward Ranking AI Overviews are generated on top of Google's live index and its ranking signals, so classic position still buys probability there, falling probability (recall the 76%-to-38% overlap collapse), but real. Practically: traditional ranking and link-authority signals keep paying on Google's AI surfaces but contribute less on ChatGPT or Perplexity, even though SEO fundamentals like crawlability and structure still matter everywhere. Budget accordingly instead of arguing about whether "authority matters" in the abstract, since it depends on the engine. ## What Actually Moves AI Citations in 2026 (and What's Folklore) Four levers have real evidence behind them: branded third-party mentions, freshness, answer-first structure, and cited evidence (sources, stats, and quotations). Two popular ones mostly don't move citations: llms.txt and schema-as-a-silver-bullet. Here's each lever, the study behind it, and how strongly it correlates, so you can spend effort where it pays.
![AI-citation signals ranked by evidence strength: branded mentions, freshness and cited sources are strongest, while domain authority is weak and llms.txt shows no measurable effect.](/blog/state-of-ai-search-2026/chart-03-what-moves-ai-citations.png)
What actually moves AI citations in 2026, ranked by evidence strength (ordinal, not a shared numeric scale). Sources: Ahrefs, AirOps, arXiv GEO study, Surfer.
SignalWhat the data showsStrengthDo this
Branded web mentions0.664 correlation with AIO visibility (Ahrefs)HighEarn third-party and community coverage before more owned content
Content freshnessPages not updated quarterly are ~3x likelier to lose citations; 83% of commercial citations from pages updated within a year (AirOps)HighPut update cycles on your cited pages, with dated changes
Sources, stats and quotationsUp to +40% generative visibility vs unoptimized content (arXiv GEO study)HighCite primary data inline; add expert quotes with attribution
Sequential heading structure2.8x citation likelihood; 87% of cited pages use a single H1 (AirOps)Medium-highOne H1, logical H2/H3 order, answer-first sections
Schema markup, 3+ types+13% citation likelihood; ~61% of cited pages use 3+ types (AirOps)Low-mediumImplement it, then stop expecting miracles from it
Domain authoritynear zero, from about -0.07 to +0.01 depending on the metric on pure LLMs (Surfer); engine-specific on AIOWeak / engine-specificStop reporting DR as an AI-visibility KPI
llms.txt~97% of files receive zero requests (Ahrefs); Google doesn't support itNo measured effectFine to ship, don't count it as work done
### The Four Levers with Real Evidence Freshness has the scariest numbers: AirOps found pages not updated quarterly are about three times more likely to lose their AI citations, over 70% of AI-cited pages were updated within the last 12 months, and 83% of citations on commercial queries come from pages updated within a year. Structure is next: sequential, logical heading hierarchy correlates with 2.8 times higher citation likelihood, 68.7% of ChatGPT-cited pages use a clean hierarchy, and 87% carry a single H1. Branded web mentions (the 0.664 correlation) you've already seen. And the [original GEO research from Princeton-affiliated authors](https://arxiv.org/abs/2311.09735) showed that adding cited sources, statistics, and quotations lifts generative-engine visibility up to 40% versus unoptimized content, the finding the entire discipline is named after. The working checklist version of all four lives in [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search). ### Does llms.txt Work? Barely, in 2026 The proposal is appealing and the adoption data isn't. Ahrefs analyzed server logs across roughly 137,000 domains and found about 97% of published llms.txt files received zero requests, from bots or humans. Google's search advocates have said plainly that Google doesn't use it. We keep a live status check in [does llms.txt work](https://geotoolbox.ai/blog/llms-txt), because this could change, but in mid-2026 llms.txt rarely does anything measurable. Ship it in an afternoon if you like; just don't report it as AI optimization. ### Schema Helps a Little Schema earned a middle verdict: pages with three or more schema types show about 13% higher citation likelihood, and 61% of cited pages use three-plus types. That's a real and modest effect, worth implementing but not a strategy on its own. The right frame for [schema markup](https://geotoolbox.ai/glossary/schema-markup) in 2026: it's connective tissue that helps engines resolve entities and structure, and the specific types worth shipping are in our guide to [schema markup for AI search](https://geotoolbox.ai/blog/schema-markup-for-ai). If a vendor pitches schema as the AI-visibility silver bullet, that 13% lift is the number to hold them to. ## The Small-Site Playbook: You Over-Index on AI Traffic If you run a small site, the data leans your way: small sites earn a higher share of AI traffic than giant ones, and AI engines routinely cite low-authority pages over big brands. You don't need a big budget. You need three moves: be reachable, be talked about, be extractable. ### Why Small Sites Over-Index on AI Traffic Ahrefs' traffic data shows sites under 10,000 monthly visits drawing about 0.3% of their traffic from AI, three times the ~0.1% share that sites with over a million visits see. Add the citation math from earlier: authority barely predicts citations on LLM engines, 26% of brands have zero AI Overview presence, and query fan-out selects precise passages over powerful domains. Classic search compounded advantages for incumbents; AI retrieval, at least in 2026, levels part of that slope. There's a geographic edge too: AI Overviews trigger most often in Indonesia, the Philippines and Mexico, so sites serving those markets may find extra opportunity there, though it needs local query and language validation. ### The First Three Moves That Aren't Hype Move one: confirm you're reachable. A surprising number of sites block GPTBot, PerplexityBot or ClaudeBot by accident, via CDN defaults or a copy-pasted robots.txt. Our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) lists every bot that matters and what to allow, and the free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) grades your site's crawlability, structure and machine-readability in one pass. Move two: earn third-party mentions where your buyers already talk. One genuine Reddit thread, one review-platform profile with real reviews, one YouTube walkthrough. That's the 85%-third-party lever scaled to a small team. Move three: restructure one flagship page answer-first. Single H1, question-shaped H2s, a 40-60 word direct answer under each, updated date. That's the 2.8x structure lever applied where it counts. Measure before and after; if you'd rather compare tooling first, we've reviewed the current field in [the best generative engine optimization tools](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools). Run the reachability check before anything else. Every other lever in this article assumes engines can fetch your pages; a blocked crawler silently zeroes out the rest of the playbook. ## Where This Leaves You The 2026 studies disagree on the decimals and agree on the direction: AI visibility is decoupling from rankings, migrating off your own domain, and concentrating in engines that don't behave alike. Every contradiction in the data resolves to a panel, a date, or a brand tier, which means the next contradictory headline you read probably isn't wrong either. It's just partial. The one thing you can act on today is measurement, because you can't fix what you can't see, and 45% of your peers can't see it. If you'd rather not spend Friday afternoons re-running prompts by hand, [geotoolbox's domain overview](https://geotoolbox.ai/features/domain-overview) runs a fixed prompt set across ChatGPT, Perplexity, Gemini, Claude, Google's AI surfaces, Microsoft Copilot and Grok, and turns mentions, citations and answer position into the trend line this article keeps asking you for. ## Frequently Asked Questions ### How much of my traffic drop is from AI Overviews vs a Google core update? Separate the timelines first: match your drop's start date against AI Overview rollout dates for your query types and against Google's confirmed core-update dates. Then check GSC: impressions holding while clicks fall points to AI Overviews; both falling together points to a ranking loss. Informational pages lose most to AI Overviews, transactional pages least. ### How do I check if ChatGPT, Perplexity or Gemini is citing my website? Build a fixed set of 10-20 real buyer questions, run each across the engines, and record whether you're mentioned, cited as a linked source, and where in the answer you appear. Repeat weekly, because answers change run to run. A single manual check is sampling noise rather than a measurement. ### Is AI search traffic worth optimizing for if it's only 0.1% of my traffic? For most sites, yes. The 0.1% to a few percent share undercounts reality because part of AI traffic hides in your Direct channel, and the visitors it does send convert several times better than classic organic. Small sites also earn roughly triple the AI-traffic share of large ones, so the smaller you are, the stronger the case. ### What's the difference between a mention and a citation in AI search? A mention is your brand named inside the answer text; a citation is your URL linked as a source. They correlate weakly, with overlap as low as 30% on Gemini, and brands earning both are about 40% likelier to stay visible. Track them as two separate metrics. ### Is AEO the same as GEO, and which should I be doing? Same discipline, different labels. GEO emphasizes generative engines, AEO emphasizes answers, LLMO emphasizes models, and all three optimize for the same outcome: being mentioned and cited in generated answers. Pick whichever term your team likes and do the work; about 80% of it is durable SEO either way. ### Does llms.txt actually do anything yet? Not measurably. Around 97% of published llms.txt files receive zero requests according to Ahrefs' log analysis, and Google has said it doesn't use the file. It costs little to publish and may age well, but in 2026 it belongs at the bottom of the list, after reachability, structure and third-party presence. ### Why does a weaker competitor get cited by AI when I don't? Usually because citation doesn't follow authority. The likeliest causes, in order: they're discussed more on third-party and community platforms, their pages answer questions in extractable answer-first passages, their entity is clearer to the engines, and their content is fresher. Domain strength correlates near zero with citations on ChatGPT and Perplexity, so "we're the bigger brand" doesn't enter the equation. ## Sources and Methodology We reconciled the major public 2026 AI-search datasets. Figures are attributed inline to their source and dated where the source provides a date, and single-source or unverified figures are flagged in the text. - Ahrefs, AI Overview citations and the rank-citation link (76% to 38% overlap, 58% CTR, 83% non-Google): ahrefs.com/blog/ai-overview-citations-top-10 - `ahrefs.com/blog/ai-overview-citations-top-10` - Ahrefs, AI SEO statistics (the 0.664 branded-mention correlation, 23x conversion, 97% zero-request llms.txt, small-site AI-traffic share, and other cited Ahrefs figures span this and the study above): ahrefs.com/blog/ai-seo-statistics - `ahrefs.com/blog/ai-seo-statistics` - Google, I/O 2026 announcement (AI Mode at 1B+ monthly users): blog.google - `blog.google/products-and-platforms/products/search/search-io-2026` - Pew Research Center, AI summaries and click behavior: pewresearch.org - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` - AirOps, The 2026 State of AI Search: airops.com/report/the-2026-state-of-ai-search - `airops.com/report/the-2026-state-of-ai-search` - Goodie, 2026 AI Search Traffic Report: higoodie.com/blog/ai-search-traffic-report-2026 - `higoodie.com/blog/ai-search-traffic-report-2026` - Semrush, 2026 AI Visibility Index: ai-visibility-index.semrush.com - `ai-visibility-index.semrush.com` - Amsive, AI Overviews click-study: amsive.com - `amsive.com/insights/seo/google-ai-overviews-new-research-reveals-how-to-navigate-click-drop-off` - Seer Interactive, AIO impact on Google CTR (2026 update): seerinteractive.com - `seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update` - Aggarwal et al., GEO: Generative Engine Optimization (arXiv): arxiv.org/abs/2311.09735 - `arxiv.org/abs/2311.09735` - Surfer, domain-authority vs AI-citation correlation study (5M citations, June 2026) - XFunnel, B2B AI-citation study (768K citations, 2026) - SE Ranking, review-platform AI-citation study (2026) - eMarketer, generative engine optimization research (2026) - BrightEdge, AI Overview prevalence tracker (2026) - Adobe, retail and travel AI-traffic growth data (Oct 2024 to May 2026), surfaced via the Semrush AI Visibility Index - Sentiment on "generative engine optimization" (~28% negative) is an informal DataForSEO content-analysis scan the author ran in July 2026, not a formal study --- ## What Is Agentic AI? A Plain-English Guide for Brands > Agentic AI explained without the hype: what it is, how AI agents work, how it differs from generative AI, and what it means for whether AI finds your brand. - Canonical: https://geotoolbox.ai/blog/agentic-ai - Published: 2026-07-05 · Updated: 2026-08-05 "Agentic AI" is the most-hyped and most-confused term in artificial intelligence right now. People use it to mean three different things, vendors promise it as magic, and most of what gets written quietly oversells it. This guide cuts through that. In plain English: what agentic AI actually is, how AI agents work, how it differs from generative AI, and the part almost nobody explains, what it means for whether AI finds and recommends your brand. The short version up front: agentic AI is real, it is genuinely useful on the right tasks, and it is much further from "set it and forget it" than the headlines suggest. ## What Is Agentic AI? **Agentic AI is artificial intelligence that can pursue a goal on its own**, breaking it into steps, using tools, taking actions, and adjusting along the way, with limited human supervision. Instead of answering one prompt at a time, an agentic system is handed an objective and works toward it. The word that matters is **agency**: the capacity to act independently and on purpose. A regular chatbot waits for you to ask. An agentic system decides what to do next. Take a simple example. Ask a normal AI assistant to "write five ad headlines" and you get five headlines. Give an agentic system the goal "improve this campaign's conversion rate this month" and it can pull the performance data, find the weak ad groups, draft variants, launch the approved ones, watch the results, and flag what it cannot decide on its own. That shift, from a tool you operate to a system that operates toward a goal, is the whole idea. It is also a question a lot of people are suddenly asking. The capability is new enough that the words are still settling, and in everyday use people often say "AI agent" and "agentic AI" loosely. A working definition that holds up across the serious sources: **agentic AI is a goal-directed system that combines reasoning, planning, memory, tool use, and action, with a varying amount of autonomy and human oversight.** One thing to fix in your head before the hype takes over: autonomy here is a dial, not a switch. Most agentic systems worth running today operate with a human watching, not hands-off, and why that matters comes later. ## Agentic AI vs Generative AI vs AI Agents Three terms get used as if they mean the same thing, and they do not. Sorting them out is the fastest way to understand the whole topic. The cleanest way to hold it in your head is a chain. **Generative AI creates. An AI agent acts. Agentic AI orchestrates.** Generative AI is the engine. It produces text, images, code, or audio in response to a prompt. It is reactive: you ask, it answers, and it does not do anything else. ChatGPT writing an email or Midjourney making an image is generative AI at work. An AI agent is that engine put to work with a body. Take a generative model, give it tools (a web browser, an API, access to your CRM) and a task, and it can take actions, not just produce words. An agent can look up a price, file a ticket, or send a draft. An agent is usually focused on one job. Agentic AI is the system around the agents. It is the goal-driven setup that plans across many steps and often coordinates several agents toward one objective. IBM frames it neatly: agentic AI is the framework, and AI agents are the building blocks inside it. Think of a smart home where one system manages your energy use by directing separate agents for the thermostat, the lights, and the appliances.
 Generative AIAI agentAgentic AI
What it doesCreates content from a promptTakes actions using toolsPursues a goal across many steps
PostureReactive (waits to be asked)Task-focused (does one job)Proactive (works toward an objective)
Needs a human toPrompt every outputSet the task and guardrailsSupervise and approve key actions
Marketing exampleDraft a landing pagePublish the page and post the linkTest variants, track results, shift budget
So is ChatGPT an agentic AI? On its own, no. The base model is generative AI. Switch on ChatGPT Work, OpenAI's agentic mode (it replaced the original "agent mode" in July 2026), and connect it to tools, and the same model starts to behave agentically. The label depends on what the system can do, not on the brand name. If you want the underlying mechanics, our glossary entry on the [AI agent](https://geotoolbox.ai/glossary/ai-agent) breaks down the parts. ## How Do AI Agents Actually Work?
![The four-step loop an AI agent runs: perceive, reason, act, learn.](/blog/agentic-ai/how-ai-agents-work-loop.png)
Every agent runs the same loop: perceive, reason and plan, act, then learn from the result.
Under the hood, an agent runs a loop. It is simpler than the marketing makes it sound. 1. **Perceive.** The agent gathers information from its environment: a user's request, data from an API, the contents of a web page or a database. 2. **Reason and plan.** A [large language model](https://geotoolbox.ai/glossary/large-language-model) interprets that information, works out what the goal needs, and decides on the next step. 3. **Act.** The agent uses a tool to do something in the real world: call an API, run a search, update a record, send a message. 4. **Learn.** It checks the result, and feeds what happened back into the next loop. The language model is the reasoning engine, but a model alone is not an agent. What turns it into one is the tools and the memory. [Anthropic draws the line](https://www.anthropic.com/engineering/building-effective-agents) between a workflow, where developers hard-code the steps in advance, and an agent, where models "dynamically direct their own processes and tool usage." The more the system decides for itself, the more agentic it is. Some tasks need only one agent. Bigger ones use several, coordinated by an orchestrator, sometimes called a conductor model, that hands subtasks to specialist agents and stitches the results together. You will also see older textbook categories, such as simple reflex agents and goal-based agents, but for a brand audience the loop above is the part that matters. In our own work, the loop is not theoretical. We run agents in parallel to research a topic, one reading competitor pages while another pulls keyword data, then bring the findings back together. The pattern is genuinely faster than doing it by hand. It is also where the limits show up, which is the next thing worth being honest about. ## What Agentic AI Looks Like in the Real World The clearest way to grasp agentic AI is to see what it actually does today, not what a vendor promises for next year. The consumer examples are the ones most people have already met. A research agent books a trip by comparing options and reserving the flight and hotel. A coding agent writes, runs, and fixes its own code. A shopping agent compares products across sites and adds the winner to a cart. A customer-service agent reads a ticket, checks an order, and issues a small refund within set limits. The major AI products now include agent modes. [Google Gemini](https://geotoolbox.ai/blog/what-is-gemini) has one that browses and acts, and OpenAI's ChatGPT Work does the same for ChatGPT. Anthropic's Claude can operate a computer and write code across a project. [Microsoft Copilot](https://geotoolbox.ai/blog/what-is-copilot), Salesforce Agentforce, and Perplexity all run task-doing agents, and platforms like [Microsoft Copilot Studio](https://geotoolbox.ai/blog/copilot-studio) let companies build their own. Even Apple has moved this way: at WWDC 2026 it announced an assistant that takes action across apps on a user's behalf, and a Passwords app that navigates websites to sign in and upgrade accounts for you. So is Siri agentic AI? Increasingly, yes. As soon as an assistant stops answering questions and starts completing multi-step tasks for you, it has crossed the line from generative to agentic. The impressive demos are real, but they are demos. The same agent that books a flight flawlessly on stage can stumble on a messy real-world account. The gap between a demo and a dependable production system is the source of most disappointment with agentic AI, and it is worth understanding before you buy. ## The Hype Problem: Agent-Washing and Why Projects Fail Agentic AI is real and it is also oversold. Both things are true, and a brand owner needs to hold them at once. Adoption is genuinely wide. A [PwC survey of US executives](https://www.pwc.com/us/en/tech-effect/ai-analytics/ai-agent-survey.html) found 79% said their companies were already using AI agents, and two-thirds of adopters reported measurable value. Andrew Ng, who helped popularize the word "agentic," now warns that the term has been grabbed by marketing departments and slapped on almost everything. That is where "agent washing" comes in: rebranding ordinary chatbots, scripts, and automation as "agents" to ride the trend. Gartner has flagged it directly and estimates that only a small fraction of the thousands of self-described agentic vendors are the real thing. The same firm [predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027), citing rising costs, unclear value, and weak controls. A [2025 MIT report](https://fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo/) found that 95% of enterprise generative AI pilots delivered little to no measurable return on the bottom line. Here is the part the failure stats usually hide: projects rarely die because the models are bad. They die on the unglamorous parts. The data is messy, the systems do not connect, nobody owns the governance, and the agent gets handed a job too broad to do reliably. For a non-technical buyer, one quick test separates a real agent from a repainted chatbot. Ask whether the system can take a goal, choose its own steps, and use tools to act across several of them. If it only follows a fixed script and answers what it is asked, it is automation with a new label, not agentic AI. That single question will save you a lot of money. ## Why Agents Aren't Magic: Reliability and Cost The single most useful thing to understand about agentic AI is that errors compound. A generative model only has to get one answer right. An agent has to get every step in a chain right, and small failure rates stack up fast. The math is blunt. An agent that is 95% reliable on a single step is only about 60% reliable across a ten-step task, because you multiply the odds at each step. Drop to 85% per step and a ten-step job succeeds about one time in five.
Accuracy per stepSuccess on a 10-step task
95%about 60%
90%about 35%
85%about 20%
This is why "just wait for a better model" is not the fix people hope for. A sharper model raises the per-step number, but a long chain still leaks. The fixes that actually work live in the design, not the model: break the job into smaller steps, let the agent verify and retry its own work, and route anything that cannot be undone to a human. Adding a check at each stage recovers much of the lost reliability, which is exactly why the systems that hold up are built around verification rather than raw model power. Let an agent run freely on cheap, reversible steps, and put a human checkpoint in front of anything it cannot take back: a payment, a deletion, an email to a customer. Full autonomy is for low-stakes tasks. Everything that touches money or your brand keeps a person in the loop. Cost behaves the same way: it grows with every step. An agent re-reads its accumulating context before each action, so a long task can burn far more tokens than a single chat, and the bill is hard to predict. The practical rule is to set hard usage caps and point agents at narrow, repeatable, high-value jobs rather than turning one loose on everything. Then there is accountability, which is not a hypothetical. When Air Canada's chatbot gave a customer wrong information about its bereavement fares, a tribunal [held the airline responsible for what its bot said](https://www.mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot) and rejected the argument that the chatbot was a separate entity. "The agent decided" is not a legal defense. Whatever your agent says or does is your company speaking. A newer risk is [prompt injection, which OWASP ranks as the number-one security threat](https://genai.owasp.org/llmrisk/llm01-prompt-injection/) for these systems: hidden instructions buried in a web page or document the agent reads can quietly redirect what it does. This shapes how we use agents in our own work, and the lesson generalizes. The checking is never left to the agent that did the job, because models are bad at catching their own mistakes: asked to review their own output, they tend to wave it through or repeat the error with more confidence. So a separate agent, with fresh context and no stake in the first answer, does the review, and we run the work past different models, including OpenAI's Codex and a few cheaper open-weight models, because each has different blind spots. We treat those open-weight models as adversarial second opinions, never as the source of truth. They are cheaper but hallucinate more than the frontier models, and their training is often months out of date, so they will confidently "correct" a true, current fact into a wrong one. Every flag gets checked against a primary source, and agreement between models is a weak signal rather than proof, since models trained on similar data can share the same blind spot. The part that actually compounds is what happens after a task. Agents do not quietly get better on their own. The improvement comes from a deliberate habit of capturing what went wrong on each run and folding it back into the rules the next run follows. That review loop, run by people, is what sharpens over time, not the model. It is the same lesson the failure stats keep pointing to: reliability is a matter of system design and discipline, not of waiting for something smarter. ## What Agentic AI Means for Your Brand's AI Visibility Here is the part the vendor explainers skip, and the part that should matter most to a brand. Agents are not just a tool you might use. They are increasingly the customer. When someone asks an assistant to "find the best CRM under $100 a month and start a trial," a [research agent does the searching](https://geotoolbox.ai/blog/how-does-ai-search-work), reads the options, and recommends one, often without the person ever visiting a website. Discovery, comparison, and sometimes the purchase itself happen inside the assistant. The brand that gets named wins. The others are invisible. This is not hypothetical: engines like Perplexity answer buying questions by naming specific products, and ChatGPT now drives product discovery and comparison (it pulled its in-chat Instant Checkout in early 2026, so the purchase itself completes on the merchant's site). This is already reshaping web traffic. Cloudflare reported in 2026 that [bots now make up about 57.5% of web requests](https://www.tomshardware.com/tech-industry/artificial-intelligence/bots-have-now-passed-human-traffic-online-cloudflare-boss-laments-says-agentic-traffic-wasnt-expected-to-eclipse-real-people-until-next-year), passing human traffic for the first time, a crossover Cloudflare attributes to the surge in AI agents browsing on behalf of assistants like ChatGPT and Gemini. Those agents are a new audience hitting your site, and a separate one from Google's crawler. Our own [list of AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) shows how many are already knocking. The uncomfortable twist for anyone who has invested in traditional SEO: ranking first on Google does not mean an agent will find or cite you. We see the gap in our own analytics. Some of our pages get picked up and cited by AI assistants while barely registering in Google's rankings for the same terms. The measurement of AI referrals is still rough, but the direction is clear: being cited by the model and ranking in search are two different games. Agents also do not browse the way people do. They favor clean, machine-readable information: clear titles, structured product attributes, specifications, and consistent facts repeated across the web. A thin description like "medium roast, caramel" gives an agent almost nothing to match against a query. Worse, many agents weigh how consistent your facts are across the web, so if your pricing or claims differ from page to page, the agent may quietly pick a competitor it trusts more. In practice that comes down to three things you control: publish clear, structured details a machine can parse; keep your facts, pricing, and claims consistent everywhere they appear; and make sure AI crawlers are not blocked from reaching you in the first place. This is the job we built geotoolbox for. Being chosen by an agent starts with two things you can actually check: whether AI crawlers and agents can [reach your site at all](https://geotoolbox.ai/blog/agent-ready-website), and whether the engines [actually cite your brand](https://geotoolbox.ai/blog/chatgpt-citations) when they answer questions in your space. That is the discipline behind [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), and it is closely tied to the rise of [agentic commerce](https://geotoolbox.ai/blog/agentic-commerce), where an agent, not a person, does much of the choosing. ## How to Start Without Getting Burned You do not need to be technical, or buy a platform, to start with agentic AI. You can begin inside a tool you already use, since both ChatGPT and Claude now run simple agents directly. The approach that works is narrow on purpose. Pick one repeatable task you understand well: routing inbound leads, drafting first replies to a common type of inquiry, turning a campaign report into a plain-language summary, or generating first-draft variants to test. Keep it small enough that you can check the output, and make sure a person approves anything the agent cannot take back, such as a send, a payment, or a deletion. Measure whether it actually saves time before you scale it to anything bigger. That is the same lesson our own agent work keeps teaching us. Start narrow, keep a human on the steps that matter, verify the output, and only widen the scope once the agent has earned it. The teams that get value from agentic AI are not the ones that hand it the most freedom. They are the ones that aim it carefully. ## The Bottom Line for Brands Agentic AI is the step from AI that answers to AI that acts. The hype is loud and the failures are real, but the direction is steady: more of what your customers do, including how they find and choose brands, will run through agents that search, compare, and decide for them. That makes one question worth asking now. When an AI agent goes looking in your category, does it find you, trust you, and recommend you, or does it quietly pick a competitor? The brands that win the agentic shift are the reachable, consistent, and citable ones. If you want to see where you stand, you can run a quick [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) to find out whether AI agents and crawlers can actually reach and read your site, then check whether the engines [cite your brand](https://geotoolbox.ai/features/citation-interceptor) when they answer questions in your space. Together they are the difference between being recommended by the agent and being invisible to it. ## Frequently Asked Questions ### Is ChatGPT an agentic AI? Not by default. The base ChatGPT model is generative AI, which means it creates content in response to a prompt. When you turn on ChatGPT Work (OpenAI's agentic mode, which replaced the original "agent mode" in July 2026) and connect it to tools, the same model can plan and act across steps, which makes that setup agentic. Whether something is agentic depends on what it can do, not on the brand name. ### What is the difference between AI agents and agentic AI? An AI agent is a single building block: a model with tools that can do a task. Agentic AI is the wider system that pursues a goal and often coordinates several agents to get there. In short, AI agents are the parts, and agentic AI is the framework that puts them to work. ### Is Siri an agentic AI? Increasingly, yes. At WWDC 2026 Apple announced a Siri that completes multi-step tasks across apps on your behalf, alongside features that navigate websites to sign in for you. Once an assistant starts completing multi-step jobs rather than only answering questions, it has crossed from generative into agentic territory. ### Is agentic AI just hype? It is both real and oversold. Adoption is genuine, but Gartner expects more than 40% of agentic projects to be cancelled by 2027, and many "agents" are rebranded chatbots, a pattern called agent washing. The technology works best on narrow, repeatable tasks with a human checking the important steps. ### What does agentic AI mean for SEO? It adds a new visibility game on top of search rankings. Agents read and recommend from structured, machine-readable information and look for consistency across the web, so ranking on Google no longer guarantees that an AI agent will find or cite you. The goal shifts from being indexed to being chosen. ### Do I need to know how to code to use agentic AI? No. You can start with simple agents directly inside tools like ChatGPT or Claude, and many no-code platforms let you assemble agents by connecting steps together. Begin with one small, reversible task before you scale up. ## Sources - Anthropic, Building Effective Agents - `anthropic.com/engineering/building-effective-agents` - PwC, AI Agent Survey - `pwc.com/us/en/tech-effect/ai-analytics/ai-agent-survey.html` - Gartner, Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027 - `gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027` - MIT NANDA via Fortune, 95% of generative AI pilots at companies are failing - `fortune.com/2025/08/18/mit-report-95-percent-generative-ai-pilots-at-companies-failing-cfo` - McCarthy Tétrault, Moffatt v. Air Canada: A Misrepresentation by an AI Chatbot - `mccarthy.ca/en/insights/blogs/techlex/moffatt-v-air-canada-misrepresentation-ai-chatbot` - OWASP, GenAI Top 10: Prompt Injection (LLM01) - `genai.owasp.org/llmrisk/llm01-prompt-injection` - Cloudflare via Tom's Hardware, Bots Have Now Passed Human Traffic Online - `tomshardware.com/tech-industry/artificial-intelligence/bots-have-now-passed-human-traffic-online-cloudflare-boss-laments-says-agentic-traffic-wasnt-expected-to-eclipse-real-people-until-next-year` --- ## Claude vs Gemini: Which Is Better? An Honest Look Under the Hood > Claude vs Gemini, compared honestly: how each is built, how they search and cite the web, who wins which task, and which one to show up in. August 2026. - Canonical: https://geotoolbox.ai/blog/claude-vs-gemini - Published: 2026-07-05 · Updated: 2026-08-14 Claude vs Gemini is usually framed as a contest with a winner. It is not one. As of July 2026 Anthropic's Claude and Google's Gemini are close enough at the frontier that the honest answer is "it depends on the task," and the two land in genuinely different places: Claude leans into writing and careful reasoning, Gemini into live search, multimodal, and the Google ecosystem it lives inside. There is also a third question almost nobody asks, and it matters most if you publish: which one cites your brand when a customer asks. This is the comparison done honestly, written for people who publish content rather than build models. New to either? Start with [what Claude AI is](https://geotoolbox.ai/blog/what-is-claude-ai) or [what Google Gemini is](https://geotoolbox.ai/blog/what-is-gemini).
![Scorecard of where Claude and Gemini each lean, from writing to Workspace integration.](/blog/claude-vs-gemini/claude-vs-gemini-scorecard.png)
Per-row leans, not a winner: Claude for depth on text and code, Gemini for reach across the live web, media, and Workspace.
## Claude vs Gemini at a Glance Here is the honest version before the detail. Both are excellent, and the real differences sit at the edges that matter to your specific work. Model versions move fast here, faster than most comparisons admit, so check the update date at the top of this page before you act on anything below.
 ClaudeGemini
MakerAnthropicGoogle (DeepMind)
Top models (July 2026)Opus 5, Sonnet 5, and the higher-tier Fable 5 (Opus 4.8 now legacy)Gemini 3.1 Pro (still in preview) and 3.6 Flash, with Gemini 3.5 Pro still in partner testing (not yet released)
Leans best atLong-document work, writing, agentic codingLive web research, multimodal work, and anything inside Google Workspace
Entry pricePro around $20/monthGoogle AI Pro around $20/month (AI Plus around $5)
Context windowUp to 1M tokens on Opus 5 and Sonnet 5, in the app as well as the API1M tokens on Gemini 3.1 Pro; Gemini 3.5 Pro is rumored at 2M (no official spec yet)
Web search and citationsWeb search built in; fewer, more authoritative citations; strong at grounding a document you give itGrounded in live Google Search; Deep Research across many sources; more real-time coverage
MultimodalText, image, and document input; no native image, video, or audio generationGenerates images, video, and audio (via Google's media models); multimodal input
Standout featureClaude Code; the careful, low-fluff characterDeep Google Workspace integration; native media generation
Free tierYes (a Sonnet-class model)Yes (the Gemini app, with limits)
Both are [large language models](https://geotoolbox.ai/glossary/large-language-model) built on the same basic idea, so a feature checklist only gets you so far. The differences that actually predict which one you will prefer come from how each was trained and how each handles the live web, which is where the real decision lives. ## Is Claude Better Than Gemini? **No, and neither is the other one.** "Better" is the wrong question, because the answer flips by task. The honest split is clear: Claude tends to win on long-form writing and careful, low-fluff reasoning, Gemini wins on live research, multimodal work, very long context, and anything that touches Google Workspace, and coding is genuinely contested. Most people who lean on either one heavily end up using both, the same pattern we found comparing [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt). That sounds like a dodge, so here is the specific shape of it. If your day is drafting, editing, and reasoning over long documents, Claude usually feels sharper and less padded. If your day involves pulling current information off the live web, generating an image or a video, working across a 2-hour meeting transcript, or living inside Gmail and Docs, Gemini's reach is hard to match because it is built into the products you already use. For code, the picture is close and depends on the job, which is a contradiction worth taking seriously rather than explaining away. There is one practical catch that the "just pick the smarter model" advice ignores, and it is the single most common complaint we see from heavy Claude users: **usage limits.** Claude's paid tiers cap how much you can send in a window more tightly than Gemini's, and people hit that wall mid-task. Gemini's limits are generally looser, and its free and low-cost tiers are more generous. A model that is slightly better per answer but stops you cold is not obviously the better tool, which is a real reason Gemini wins for some people who never argue it is the smarter model. So the useful question is not which one is better. It is better at what, for whom, and measured how, which is what the rest of this comparison gets specific about. ## How Claude and Gemini Actually Work (and Why You Feel the Difference) **Under the hood, both run the same kind of engine and then diverge in two ways that explain most of how they behave.** Each is a transformer that predicts the next token, trained on a large slice of the internet. The character you talk to comes from a second stage, alignment, and from the product wrapped around the model. Both of those diverge sharply between Claude and Gemini, and that is the part most comparisons skip. The training recipe differs. Claude's behavior is shaped by reinforcement learning from human feedback plus [Constitutional AI](https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback): the model first critiques and revises its own answers against a written set of principles, and the reinforcement step is then driven by preferences an AI generates against those same principles, rather than by human ratings alone. Gemini comes out of Google DeepMind with a different emphasis: it was [built to be natively multimodal](https://deepmind.google/models/gemini/), trained on text, images, audio, and video together rather than having vision bolted on afterward, which helps how it works across media. Different recipes, different defaults. We cannot inspect either model from the outside, so treat any claim about why it behaves a certain way as an informed read, not a readout. The deeper mechanism lives in our companion pieces on [how Claude works under the hood](https://geotoolbox.ai/blog/how-does-claude-work) and [how ChatGPT actually works](https://geotoolbox.ai/blog/how-does-chatgpt-work), whose mechanics apply to any transformer. The second divergence is the product, and it matters more than the model for most daily decisions.
 The modelThe product
What it isA trained network that turns a prompt into more textThe app around it: tools, memory, web search, media generation, integrations, safety, UI
ClaudeOpus 5 / Sonnet 5 / Fable 5, shaped by Constitutional AIclaude.ai and Claude Code: web search, file handling, artifacts, agentic coding
GeminiGemini 3.1 Pro / 3.6 Flash (multimodal input, text output)The Gemini app and Workspace: live Google Search, image and video generation via Google's media models, Gmail and Docs, Deep Research
When people say "Gemini can make a video" or "Gemini is right there in my Gmail," they are usually describing the **product**, not the raw model. Keeping the two apart is the difference between a comparison that holds up and one that ages badly the next time either company ships a feature. A capability gap today is often a product decision, not a permanent limit of the underlying model. ## How Each One Searches the Web and Cites Sources **This is where the two pull furthest apart.** A model only knows the world up to its training cutoff. As of July 2026 that is around May 2026 for [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5), per Anthropic's docs, four months fresher than the legacy Opus 4.8 and Sonnet 5, which both stop around January 2026. Anything newer, including most recent facts about your company, reaches the model only at query time through live web search, the context you paste in, or files you upload. Gemini's structural advantage here is obvious once you say it out loud: it is made by the company behind Google Search. Its answers can be [grounded in live Google Search](https://ai.google.dev/gemini-api/docs/google-search), and its Deep Research mode fans a question out across many sources and synthesizes them into a cited report. Tools like NotebookLM extend the same reach to your own documents. For anything that depends on current information, that reach is a real edge, and the citations come with clickable links you can check. Claude's web search is newer and tends to be more conservative, pulling fewer sources and favoring authoritative or technical ones, and it is unusually strong when you paste a long document into its [context window](https://geotoolbox.ai/glossary/context-window) and ask it to reason over that instead of the open web. If your question is "summarize and pressure-test this 80-page contract," Claude's grounding on the material you hand it is a different and often better job than searching the web at all. Now the part both share, and the one to take seriously. When web search is **off**, each model answers from frozen training memory, and either one can produce a confident, authentic-looking citation that does not exist. This is not lying. It is a known failure mode where the model recognizes the shape of a source and fills in plausible details, which is why both tools invent journal articles and URLs that were never real. We cover the mechanism in [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations). Turning on web search reduces it sharply but does not make it zero, so the citation a model hands you is a lead to verify, never a guarantee. Both sit on the same underlying machinery of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work); they just feed it from different indexes. ## Claude vs Gemini for Writing, Coding, and Research **Sorted by job, the picture gets clearer than any overall winner.** Here is where each one tends to land, with the caveat that matters for each.
Use caseUsually leansWhyThe caveat
Long-form writing and editingClaudeMore natural prose, less filler, holds a long document in working memory wellLeans agreeable; you have to ask it directly for hard critique
CodingContested, slight Claude leanStrong agentic coding and Claude Code, plus a coding-benchmark edgeA benchmark lead is not an implementation lead; Gemini competes hard on cost and context
Live research and current infoGeminiLive Google grounding and Deep Research pull more, fresher sourcesMore sources is not more accuracy; the summary can rest on weak ones
Generating images, video, and voiceGeminiNative generation across image, video, and audioClaude does not generate media at all; for created images or video, Gemini is the tool
Very long documentsSplitBoth run near a 1M-token window today, with 3.5 Pro rumored to go higher; Claude tends to stay coherent across a long oneAdvertised context is not the same as reliable recall, so test it on your own documents
Google Workspace tasksGeminiBuilt directly into Gmail, Docs, and SheetsNative beats a connector; Claude reaches these through add-ons, not from inside
On **writing**, Claude's edge is real but comes with a tell: it leans agreeable. Ask it to critique your draft and it will often soften the verdict, so you have to explicitly tell it to be harsh. Gemini's prose is strong and stylistically more neutral; some readers prefer that, others find Claude's voice less corporate, and it is a matter of taste more than a quality gap. Neither replaces an editor. On **coding**, Claude currently leads the headline benchmark, SWE-bench Verified, though that score is self-reported and the lead has changed hands before. The real-world choice is closer and depends on your stack: Claude has the stronger reputation for long, multi-file agentic work, while plenty of developers get equal or better results from Gemini on their own code, and Gemini runs far cheaper at volume. Benchmark leadership and the build that actually compiles on your machine are different claims. On **research**, Gemini's live grounding is a clear advantage for anything time-sensitive, but the same caution from our [comparison of ChatGPT vs Perplexity](https://geotoolbox.ai/blog/chatgpt-vs-perplexity) applies: a research-style answer is only as good as the sources it grounded on. Read the links, do not trust the summary. One nuance the tables flatten: both models read images, screenshots, and PDFs you hand them well, so on understanding multimodal input they are close. The gap is generation, where only Gemini makes new images, video, and audio. ## Claude Code vs Gemini CLI: Coding in the Terminal **The most useful coding comparison right now is not the chat models, it is the terminal agents, and that is the matchup nobody covers.** Both companies ship a command-line tool that turns the model into an agent that reads your codebase, plans, edits multiple files, runs commands, and checks its own work. They take different bets. Claude Code is the more mature of the two on long, multi-step tasks. It holds intent across a complex job, recovers from its own errors reasonably well, and connects to external tools through the Model Context Protocol. Its main cost is exactly that: cost, plus the usage caps that bite hardest during heavy agentic sessions. [Gemini CLI](https://github.com/google-gemini/gemini-cli) is open source, and Google has paired it with an unusually generous free quota and a 1M-token context window. Running on Gemini's Flash-class models, such as [Gemini 3.6 Flash, the current default](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/) and tuned for sustained agentic and coding work, it is fast and cheap enough to leave running, and it slots into Google's broader developer tooling. For large repositories where context size matters, or for anyone watching spend, it is a serious option rather than an also-ran. The honest call: Claude Code still feels a step ahead on agentic coherence for hard, long tasks, while Gemini CLI wins on cost, openness, and raw context. Pick by your budget and which ecosystem you already live in, and re-test after any major release, because this is the fastest-moving corner of the whole comparison. ## What the Benchmarks Actually Say (Read Them Skeptically) **Benchmarks are the most quoted and least understood part of any comparison.** They are useful for spotting which models are roughly in the frontier tier and nearly useless for picking a daily driver. Here is the honest state of the major ones as of July 2026.
BenchmarkWhat it measuresReported leader (July 2026)Read it with
SWE-bench VerifiedResolving real GitHub issues (coding)Claude's frontier models lead: Opus 4.8 at 88.6% (the model these scores were run against, since succeeded by Opus 5) and Fable 5 at 95%, against 80.6% for Gemini 3.1 Pro, roughly an eight-point gap, wider than most comparisons implyScores are largely self-reported and swing with the test harness
GPQA DiamondGraduate-level science questions (reasoning)Contested; the two trade the leadScores sit near the ceiling, so tiny gaps look bigger than they are
Humanity's Last ExamHard expert reasoning, no toolsBoth rank at the frontier; the lead has traded between themA young, brutal benchmark; low scores with wide error bars
LMArenaHuman head-to-head preference votesClaude and Gemini both rank among the top models in textThis measures what people prefer, not what is correct
Long-context recallFinding facts buried in a huge inputGemini's larger window helps on the biggest inputsRecall still degrades in the middle of long contexts for every model
Four things keep these numbers from meaning what they look like. First, most headline coding scores are **self-reported** by the labs, and the same model can swing several points depending on the scaffolding around it. Second, the older general-knowledge tests are effectively saturated, with everyone scoring in the high 80s and 90s, so they no longer separate the top models. Third, the leads flip almost monthly as each lab ships, so any single number is a snapshot, not a standing. Fourth, [LMArena](https://arena.ai/leaderboard) rewards the answers people prefer, which tracks quality but also rewards confident, well-formatted responses that are not always right. The practical takeaway is dull but true: by mid-2026 the frontier models from both labs are close enough that benchmark gaps rarely decide a real workflow. The benchmark that counts is your own work, run through both. ## Pricing, Usage Limits, Privacy, and the Google Ecosystem **The specs people compare are rarely the things that decide it. Four practical factors matter more.** **Price and limits.** At the consumer level the two are close: Claude Pro and [Google AI Pro](https://one.google.com/about/google-ai-plans/) both sit around $20 a month, with higher tiers (Claude Max and Google AI Ultra) running $100 to $200. Gemini has the cheaper entry point, an [AI Plus tier around $5](https://geotoolbox.ai/blog/gemini-pricing) (cut from $8 in June 2026), and generally looser usage caps. That last part is the one heavy users feel most: Claude's tighter limits are the friction they hit soonest, while Gemini lets you keep going, and Gemini's free tier is the more generous of the two. On the API, Gemini's price edge is real but narrower than the shorthand suggests: its Pro tier undercuts Claude's Opus per token ($2/$12 against $5/$25). But "Flash is cheaper than almost anything" no longer holds: Gemini 3.6 Flash runs $1.50/$7.50 per million tokens, while Claude's Haiku 4.5 is $1.00/$5.00, cheaper on both. Claude's Sonnet 5 even undercuts Gemini 3.1 Pro on output at $2/$10 (a rate Anthropic made permanent in August 2026). At volume, which one is cheaper depends on the pair you compare, not the logo. **The Google ecosystem.** This is Gemini's quiet trump card. It is built into Gmail, Docs, Sheets, and Android, so for the millions of people already living in Workspace, Gemini is the assistant that is simply there, in the document, with the context already loaded. Claude reaches those tools through add-ons, which is not the same as being native. If your work happens inside Google's products, that integration can outweigh any model-quality difference. **Privacy.** As of July 2026 both makers train on consumer chats by default and let you opt out, while business, enterprise, and API tiers are excluded from training by default. Anthropic changed Claude's consumer default in 2025, so whatever you remember about Claude not training on your data may be out of date. Some people are wary of Gemini specifically because it is Google, the same company whose business runs on data and ads; that is fair to weigh, though the practical training controls are similar on both sides. If your work is sensitive, use a business tier or check the data settings on whichever one you run. **Drift.** Both models change under you. New versions ship constantly, old ones retire, and people regularly report a model feeling worse after an update, sometimes real, sometimes perception. Re-test your own workflow after any major release rather than trusting last quarter's verdict. ## Which One Should Your Brand Appear In? **If you publish or market anything, the question that pays is not which tool you use. It is which one names your brand when your customer asks.** Both Claude and Gemini are answer engines now. People ask them which product to buy, which agency to hire, which tool is best, and the answer either includes you or it does not. In that moment the citation is the impression, and the ones you lose stay invisible unless you go looking. This is not a fringe behavior, and the clearest measured case so far is Google's own search. A 2025 [Pew Research Center study](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) found that when an AI summary appeared in Google results, users clicked a traditional result link in just 8% of visits, against 15% when there was no summary, nearly halving the click-through. That study measures Google AI Overviews and U.S. Google users, not Claude or Gemini directly, but it captures the same shift: as more answers get consumed without a click, being inside the answer matters more than ranking below it. Here the two engines split in a way that is specific to this matchup. Gemini draws on Google's own index, so your classic Google SEO and your Gemini visibility overlap heavily, which is convenient if you already rank and a problem if you do not. The overlap is not total, though: Google's AI surfaces use their own retrieval signals, so ranking first does not guarantee a citation. Claude pulls from its own newer web search, a more independent surface that leans toward authoritative sources. Showing up in one says little about the other, which turns into two concrete jobs. For Gemini and [Google's AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo), the work is mostly strong classic ranking, clean crawlable HTML, and structured data. For Claude, [getting cited](https://geotoolbox.ai/blog/claude-seo) rewards plain, well-organized, authoritative prose over keyword-padded pages. You cannot control which model a customer opens, and you cannot control the dice on any single answer. In our experience auditing brands across these engines, the most common and most fixable problem is simply not knowing the score: companies are absent from AI answers about their own category and have no idea, because they are still watching blue-link rankings. You cannot improve a number you are not measuring, which is the whole case for tracking your [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) per engine rather than guessing. ## How We Use Both, and Why the Loop Beats the Model A note on where this comes from. We run an SEO agency, build a GEO tool, and do security research, and most of that work runs through Claude in Claude Code. But Gemini earns a real place in the stack: when we need something current off the live web, a fast pass over a giant document, or media we cannot make in Claude, Gemini is the right tool, and we reach for it without ceremony. So this is not a spec-sheet comparison. It is how we use both every day. The thing we learned the hard way is that the model matters less than the loop around it. A model reviewing its own work is notoriously bad at catching its own mistakes. It rereads its error and confidently approves it, because the same blind spot that produced the mistake is doing the checking. The fix is not a smarter model. It is a second, adversarial pass, ideally from a different model family, because a different model fails in different places and sees what the first one cannot. So we never ship a first draft. Whether it is an article, a line of production code, or a security finding, it goes through several harsh, independent reviews before it counts. This article is an example: one model wrote it, then independent reviewers and a separate cross-model pass tore into it and caught real problems the draft was blind to. We use these models against each other on purpose, which is the same reason geotoolbox measures AI visibility by sampling many times instead of trusting one check. ## So, Claude or Gemini? How to Decide Pick by task, not by leaderboard. Reach for Claude when you are writing, editing, or reasoning over long documents, and when you want a careful, low-fluff style. Reach for Gemini when you need current information off the live web, images or video, deep Google Workspace integration, or the broader Google ecosystem. Both have free tiers, so the cheapest way to settle the debate for your own work is to put the same real task through each and stop looking for a universal winner that does not exist. For a business, the calculus is different, and it has nothing to do with which model is smarter. Your customers use both, so the question that actually moves revenue is not which one you prefer, but which one cites you when someone asks about your category, and how that compares to your competitors. That is the gap geotoolbox closes. We track what Claude, Gemini, and the other engines actually say and cite about your brand, per engine and over time, so you can see where you show up, where a competitor owns the answer, and what to fix. You can start with a free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether these engines can even read your site, then watch your presence across them in the [domain overview](https://geotoolbox.ai/features/domain-overview). The models will keep trading the lead. Whether they mention you is the part you can work on. ## Frequently Asked Questions ### Is Claude better than Gemini? It depends on the task, and that is the honest answer rather than a dodge. As of July 2026, Claude tends to win on long-form writing, document reasoning, and careful, concise output, while Gemini wins on live web research, multimodal generation, very long context, and deep Google Workspace integration. Coding is close. For most people there is no single winner, which is why heavy users often use both and switch by task. ### Is Gemini or Claude better for coding? The two are close, with a slight edge to Claude on hard, multi-step agentic work, and a real cost-and-context edge to Gemini. Claude Code has the stronger reputation for holding a long task together, while Gemini CLI is open source, cheaper to run, and ships a 1M-token context window that helps on big repositories. Benchmark results move with each release, so the only test that settles it is your own codebase. ### Does Gemini have a bigger context window than Claude? Right now they are close, and the answer depends on the model. On paid plans Claude gives you 1M tokens when chatting with Opus 5 or Sonnet 5, and the same 1M through the API. Gemini's app gives 1M tokens to AI Pro and Ultra subscribers, 128K on AI Plus and 32K without a subscription. Gemini 3.5 Pro is rumored to target 2M, though Google has published no official spec; if true, that would pull Gemini ahead once it ships. A bigger window is not the same as better recall, though: every model gets less reliable at finding facts buried in the middle of a very long input, so window size is a ceiling, not a guarantee. ### Is Claude or Gemini cheaper? At the consumer level they are close, both around $20 a month for the main paid plan, with Gemini offering a cheaper entry tier around $5. At the API it depends on the pair: Gemini 3.1 Pro ($2/$12 per million tokens) undercuts Claude's Opus 5 ($5/$25), but Claude's Haiku 4.5 ($1/$5) undercuts Gemini's 3.6 Flash ($1.50/$7.50) on both input and output. Claude's tighter usage limits also cut heavy sessions off sooner. Which one bills less at volume depends on which models you actually run, not on the logo. ### Which one hallucinates less, or cites sources better? Gemini's live grounding in Google Search gives it more current sourcing with clickable links, while Claude tends to pull fewer sources, favoring authoritative ones, and is strong at grounding answers in a document you provide. With web search off, both answer from training memory and either can invent a confident, fake-but-plausible citation. The safe rule is the same for both: treat every cited source as a lead to verify, not proof. ### Is Claude or Gemini better than ChatGPT? No model is the single most powerful; at the frontier, Claude, Gemini, and ChatGPT trade the lead. ChatGPT is the broad generalist with the widest tool ecosystem and native image generation, Gemini owns live Google grounding and Workspace integration, and Claude leads on long-form writing and agentic coding. The gemini vs claude decision usually comes down to ecosystem and task; if you are weighing the wider field, our [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt), [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt), [Grok vs Gemini](https://geotoolbox.ai/blog/grok-vs-gemini), and [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude) comparisons cover those pairings directly. ### Does it matter which one my brand shows up in? Yes, because your customers use both, and the two engines draw on different sources, so being mentioned in one does not mean being mentioned in the other. Gemini leans on Google's index, so it tracks your classic SEO; Claude uses its own web search and rewards authoritative, well-structured pages. As people get more answers without clicking, the practical step is to measure how often each engine cites you, per engine and over time, and fix the gaps rather than guess. ## Sources - Anthropic - Constitutional AI: Harmlessness from AI Feedback - `anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback` - Anthropic - Claude models overview (current models, context windows, pricing) - `platform.claude.com/docs/en/about-claude/models/overview` - Google DeepMind - Gemini models - `deepmind.google/models/gemini` - Google - Gemini 3.5: frontier intelligence with action - `blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5` - Google - Grounding with Google Search (Gemini API) - `ai.google.dev/gemini-api/docs/google-search` - Google One - AI plans (Free, AI Plus, AI Pro, AI Ultra) - `one.google.com/about/google-ai-plans` - Google - Gemini CLI (open-source terminal agent) - `github.com/google-gemini/gemini-cli` - SWE-bench - software engineering benchmark - `swebench.com` - LMArena - human-preference LLM leaderboard - `arena.ai/leaderboard` - Pew Research Center - Google users are less likely to click links when an AI summary appears - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` --- ## Microsoft Copilot Pricing in 2026: Plans, Free Tier & Cost > Microsoft Copilot pricing in 2026: the free tier, Microsoft 365 Premium, the $18/$21/$30 business plans, the base-license trap, plus Studio and Cowork credits. - Canonical: https://geotoolbox.ai/blog/copilot-pricing - Published: 2026-07-05 · Updated: 2026-08-10 Microsoft Copilot starts free, and for most personal use that is the end of the story. If you want more, the consumer paid plans run from $9.99 a month, with Microsoft 365 Premium at $19.99 putting Copilot inside your Office apps. For work, the business plans run from $18 to $30 per user per month. As of June 2026, those are the headline numbers. The catch is the one almost every guide skips: the business plans are an **add-on**. You cannot buy Microsoft 365 Copilot on its own. You must already hold a qualifying Microsoft 365 license first, so the real cost is often two to three times the sticker. Below is every current Microsoft Copilot price, reconciled from each vendor's own pricing page, plus the only question the numbers exist to answer: which plan, if any, you should actually pay for. One disambiguation first, because the search results mix them: this is about **Microsoft Copilot**, the AI assistant. It is not GitHub Copilot, the separate coding tool for developers, which is sold on its own plans and not covered here. Microsoft also prices Security Copilot, Sales and Service Copilot, and Copilot in Dynamics 365 separately; this guide covers the everyday assistant most people mean. ## How Much Does Microsoft Copilot Cost? Every Plan at a Glance Here is the whole family in one place, at US prices as of June 2026.
PlanWho it's forPrice (US, 2026)Base license needed
Free CopilotAnyone$0None
Microsoft 365 PersonalOne person$9.99/mo or $99.99/yrNone (it is the subscription)
Microsoft 365 Family1 to 6 people$12.99/mo or $129.99/yrNone
Microsoft 365 PremiumIndividuals who want the most Copilot$19.99/mo or $199.99/yrNone (replaced Copilot Pro)
Microsoft 365 Copilot ChatWork and schoolFreeAn eligible Microsoft 365 plan
Microsoft 365 Copilot BusinessTeams up to 300 users$18 to $21/user/mo (+ base)A qualifying Microsoft 365 Business plan
Microsoft 365 Copilot (Enterprise)Larger organizations$30/user/mo (+ base)Microsoft 365 E3 or E5
Copilot StudioBuilding custom agents$200/pack/mo or $0.01/creditVaries
Copilot CoworkRunning autonomous agentsUsage-based creditsA Microsoft 365 Copilot license
Two things drive the confusion. First, **"Copilot" is not one product** but a family: a free consumer chatbot, paid consumer plans, a free work chat tier, paid business licenses, and a platform for building agents (our guide to [what Microsoft Copilot is](https://geotoolbox.ai/blog/what-is-copilot) sorts them out). Second, **the business tiers are add-ons**, which is where the budget surprises start. ## Is Microsoft Copilot Free? What the Free Tier Actually Gives You Yes, and there are two different free tiers, which trips people up. The first is **consumer free Copilot** at copilot.microsoft.com and in the mobile apps. It gives you chat, web answers with clickable citations, image generation, and voice, with usage limits during busy periods. No subscription required. The second is **Microsoft 365 Copilot Chat**, the free tier for work. If your organization has an eligible Microsoft 365 subscription, anyone signed in with a work account gets secure, web-grounded AI chat at no additional cost, with commercial data protection. The catch is in what they cannot do. Both free tiers answer from the public web and from files you upload in the moment. Neither one reaches into your own emails, documents, or meetings on its own. The thing most people want from Copilot at work, summarizing a client thread or drafting a reply from your real inbox, only comes with a paid Microsoft 365 Copilot license. If a colleague says Copilot summarized their inbox and yours will not, the difference is a paid license, not a setting. There is also a trap for admins: the free Copilot Chat lets anyone build agents grounded in work data, but running them against your organization's data is metered in Copilot Credits, not included, so a free agent can quietly generate a bill once it answers from your SharePoint and Teams content. For light personal use, the free Copilot is genuinely enough. You start paying for one of two very different things: Copilot inside your Office apps, or Copilot that reads your company's work. ## Microsoft Copilot for Individuals: Personal, Family, and Premium If you want Copilot working inside Word, Excel, PowerPoint, and Outlook for personal use, it comes bundled into a consumer Microsoft 365 subscription. There is no separate consumer "Copilot" line item anymore.
PlanPrice (mo / yr)PeopleCloud storageCopilot usage
Microsoft 365 Personal$9.99 / $99.9911 TBHigher than free
Microsoft 365 Family$12.99 / $129.991 to 66 TB (1 TB each)Higher than free
Microsoft 365 Premium$19.99 / $199.991 to 66 TB (1 TB each)Highest
Yearly billing saves up to 17 percent over monthly, and all three offer a one-month free trial. One quirk worth knowing: the Copilot AI features are tied to the **subscription owner** and cannot be shared across the household, even on Family and Premium. ### If You're Looking for Copilot Pro You cannot buy it anymore. Microsoft **stopped selling the standalone Copilot Pro subscription** (the old $20 a month consumer plan) to new customers on October 1, 2025, and folded its features into [Microsoft 365 Premium](https://www.microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse/) at $19.99 a month. Existing Copilot Pro subscriptions ran until support ended on August 1, 2026, and Microsoft's consumer pricing page now lists only Personal, Family, and Premium. This matters because **"Copilot Pro at $20 a month" is the single most common out-of-date Copilot price** still floating around. Wikipedia, older pricing guides, and even Google's AI Overview keep repeating it. And to answer the question this retirement tends to raise: no, Microsoft Copilot is not shutting down. Microsoft is renaming and expanding the lineup, not ending it. When Microsoft added Copilot to the standard Personal and Family plans, it raised those subscription prices, and not every customer wanted to pay for AI. For those who do not, a no-Copilot **"Classic"** version of Personal and Family still renews at the older, lower price, though as of 2026 it is not openly listed and usually surfaces only as a downgrade option during cancellation. ## Microsoft 365 Copilot for Business: The $18, $21, and $30 Plans This is where the per-user pricing lives, and where most of the confusion starts. There are three things to know. **Microsoft 365 Copilot Chat** is the free work tier described above: secure web-grounded chat included with an eligible Microsoft 365 subscription, plus agents billed on a metered basis. No per-user fee. **Microsoft 365 Copilot Business** is the paid add-on for smaller organizations, up to 300 users. Microsoft currently lists it at **$18 per user per month** on an annual commitment, a promotional rate reduced from the $21 list price, or $25.20 per user per month month-to-month (the $18 rate needs the annual commitment). It adds Copilot inside Teams, Outlook, Word, PowerPoint, and Excel, grounded in your work data, plus the ability to build agents in Copilot Studio. **Microsoft 365 Copilot**, the enterprise tier, is **$30 per user per month, paid yearly**. It is the full assistant for larger organizations, layered on an enterprise base license. Which tier you can buy comes down to your base license, not a headcount Copilot checks. Copilot Business attaches to a [Microsoft 365 Business plan](https://www.microsoft.com/en-us/microsoft-365-copilot/pricing), itself capped at 300 seats; the $30 Microsoft 365 Copilot attaches to enterprise E3 or E5, with no cap. So a small company on E5 can buy the $30 tier, while one over 300 users cannot use Copilot Business, in both cases because of the base plan. And the cheap headline assumes an annual commitment, so you cannot quietly trial a month at the advertised rate. ## The Real Cost: Copilot Is an Add-On You Buy on Top of a License This is where the sticker shock comes from. Microsoft 365 Copilot is **never sold on its own**. The $18, $21, or $30 buys you the Copilot layer only. To use it, every user must already hold a qualifying Microsoft 365 base license, and that base often costs more than the Copilot add-on itself. So your true monthly cost is the base plan plus Copilot. If you do not have a base license, the add-on will not even appear as purchasable in your admin center, which is why the real per-seat cost lands well above the advertised number.
Base planBase price+ CopilotTrue all-in (per user/mo)
Business Standard$14$18$32 (or a $23.50 bundle)
Business Premium$22$18$40 (or a $32 bundle)
Enterprise E3$39$30$69
Enterprise E5$60$30$90
![Microsoft Copilot true all-in cost: a required Microsoft 365 base license plus the Copilot add-on, August 2026.](/blog/copilot-pricing/copilot-true-all-in-cost.png)
The real per-seat cost is the required base license plus the Copilot add-on.
Microsoft also sells bundled "with Copilot" plans for small businesses that fold the base and the add-on into one number, like Business Standard with Copilot at $23.50 and Business Premium with Copilot at $32, which usually beats buying the pieces separately because Microsoft prices the bundle below the sum of the parts. At the top of the enterprise stack, the [Microsoft 365 E7 plan](https://www.microsoft.com/en-us/microsoft-365/enterprise/microsoft365-plans-and-pricing) at $99 per user per month is the one enterprise tier that bundles Copilot in rather than charging for it as an add-on. Now multiply by your team. Ten seats of Copilot Business on Business Standard is not $180 a month, it is $320 a month all-in, or roughly $3,840 a year. At E3, ten enterprise seats run $690 a month. It scales fast, so blanket-licensing everyone on day one is the most expensive way to start. The cheaper path is the one Microsoft's own ROI numbers depend on. [Velosio's read of Microsoft's Forrester study](https://www.velosio.com/blog/m365-copilot-pricing-calculator/) cites a 4.5x three-year return and about 1.5 hours a week freed per person, but only when adoption happens. In practice that means licensing a small pilot cohort, proving the value, then scaling, rather than paying for seats that go cold by month three. ## Copilot Studio and Cowork: The Pay-as-You-Go Credit Model There is a third billing model, and it is the newest source of bill shock: **consumption credits**, where one Copilot Credit costs about a cent. [Copilot Studio](https://geotoolbox.ai/blog/copilot-studio), the low-code tool for building custom agents, is sold as a $200 per month capacity pack that includes 25,000 Copilot Credits, or pay-as-you-go at $0.01 per credit through an Azure subscription. The trap is that different actions burn wildly different amounts: a scripted answer might cost a single credit, while a token-heavy reasoning response can cost a hundred or more. The same agent can run $8 a month or several hundred depending on how it is built and how often it runs, and capacity-pack credits expire monthly if unused. We break the credit math down in our [Copilot Studio guide](https://geotoolbox.ai/blog/copilot-studio). Then there is **Copilot Cowork**, which became [generally available on June 16, 2026](https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/) and caused the loudest pricing reaction of the year. Cowork is the [agentic](https://geotoolbox.ai/blog/agentic-ai) layer that takes a long, multi-step job and runs it in the background. Two things about its pricing surprised people. It is not a flat seat fee, and it is **not included** in the Microsoft 365 Copilot license you already pay for. Every Cowork run is billed on usage-based Copilot Credits, on top of a Copilot license, with the per-task cost depending on the model it uses (it launched on Anthropic's Claude Opus 4.8 and Sonnet 4.6 (Sonnet 5 rolled in later in June 2026), which remain available, though OpenAI's GPT-5.6 became its preferred model on July 9, 2026), the context it pulls in, and how long it runs. Because Cowork meters by the task rather than the seat, the same automation can cost pennies or several dollars per run, and you only see the bill after it finishes. If you turn it on, use the admin spending caps and budget alerts, and pilot with a real workload before assuming a fixed monthly figure. Variable AI billing is genuinely hard to forecast, which is why early users on Reddit dubbed it a new microtransaction era for productivity software. For most buyers, Studio and Cowork are not part of the decision at all. They matter only if you are building or running agents, and if you are, the rule is the same: set caps, estimate against a real task, and never assume a credit pack is a fixed cost. ## What Changed on July 1, 2026, and Other Moving Targets Copilot pricing has a few dates worth knowing, because the number you see today is not necessarily the number you paid last month. The most important nuance: the [July 1, 2026 price increase](https://www.microsoft.com/en-us/microsoft-365/blog/2025/12/04/advancing-microsoft-365-new-capabilities-and-pricing-update/) hit the **base Microsoft 365 commercial licenses, not the Copilot add-on**. Business Basic moved from $6 to $7, Business Standard from $12.50 to $14, E3 from $36 to $39, and E5 from $57 to $60. Business Premium held at $22, and the Copilot add-on price itself did not rise. Because your all-in cost is base plus Copilot, a higher base still raises what you pay overall, so the base went up, not Copilot. The promotional rates have their own clocks. Enterprise volume discounts on the $30 add-on have been reported by Microsoft partners as a ladder running from 15 percent off at 10 or more seats up to 40 percent at 1,000, expiring June 30, 2026; Microsoft does not publish that ladder itself, so treat the specifics as partner-channel reporting rather than list pricing, and budget the full $30 unless your reseller quotes otherwise. What Microsoft does document is a separate channel offer: 15 percent off for customers committing to a three-year agreement at 300-plus licenses, available June 1 through September 30, 2026. The $18 Copilot Business add-on rate is promotional: Microsoft's pricing page lists the discount as "available between July 1, 2026, and September 30, 2026," applying to the first year only. Treat $18 as a limited-time rate, not a permanent price, since it can step toward the $21 list once that window closes. The bundled "with Copilot" plans are a different story and are no longer a promotion at all — as of July 1, 2026 Microsoft 365 Business Standard with Copilot and Business Premium with Copilot became **permanent SKUs**, a change Microsoft framed as removing "the friction of selling Copilot as an add-on or as a short-term promotional offer." So $23.50 and $32 are standing prices, not a countdown. Finally, every figure here is US list pricing. The 2026 changes apply globally with local market adjustments, so the actual cost in your currency, agreement type, and purchasing channel can differ, and nonprofit, education, and government (GCC) tenants have their own Copilot pricing. International and sector buyers should confirm against their own Microsoft pricing rather than converting the US number. ## Is Microsoft Copilot Worth It? Copilot vs ChatGPT, Claude, and Gemini on Price On the consumer side, Copilot is priced in line with its rivals. Microsoft 365 Premium at $19.99 sits next to ChatGPT Plus, [Claude Pro](https://claude.com/pricing), and [Gemini AI Pro](https://gemini.google/subscriptions/), which all cluster around $20, except Premium also bundles the Office apps and 6 TB of storage. You are not paying $20 for a chatbot alone.
PlanPrice (US/mo)What you're really paying for
Free Copilot$0Web-grounded chat, image generation, voice
Microsoft 365 Premium$19.99Copilot in your Office apps, plus Office and 6 TB storage
ChatGPT Plus~$20A faster standalone chatbot many find sharper
Claude Pro~$20A standalone chatbot strong at writing and reasoning
Google AI Pro (Gemini)$19.99Gemini across Google apps and the web
Microsoft 365 Copilot$30 + baseAn assistant grounded in your company's email, files, and meetings
You do not pay the Copilot premium for a better general chatbot. For pure standalone chat, many people find ChatGPT a little sharper, and at $20 with no prerequisite license it is the simpler buy. What the $30 Microsoft 365 Copilot does that a $20 chatbot cannot is **work grounded in your own tenant**, summarizing your real meetings and drafting from your real inbox, and acting inside Word, Excel, and Outlook. That is a different category, not a more expensive version of the same thing. We go deeper in [Microsoft Copilot vs ChatGPT](https://geotoolbox.ai/blog/microsoft-copilot-vs-chatgpt), and our [Gemini pricing breakdown](https://geotoolbox.ai/blog/gemini-pricing) runs the same exercise for Google. So whether Copilot is worth it depends on the buyer: for someone who lives in Office, Premium at $19.99 is an easy yes; for a business, the real question is the all-in cost and whether the seats get used. ## Which Microsoft Copilot Plan Should You Actually Pay For? Most people overbuy. Here is how to decide, from the bottom up. **Stay free** if you mostly ask Copilot general questions and want web answers. The free Copilot covers it, and you will know you have outgrown it the day you keep wishing it could read your own files. **Pay $19.99 for Microsoft 365 Premium** if you want Copilot working inside your personal Word, Excel, Outlook, and PowerPoint, and you would value the Office apps and storage anyway. It is the cleanest consumer choice and it replaced Copilot Pro. **Pay for Microsoft 365 Copilot Business** ($18 to $21 per user, plus a qualifying base plan) if a small team needs Copilot grounded in your shared work. There is no broad free trial for the paid add-on, so use the free Copilot Chat tier to evaluate, pilot a handful of seats month-to-month to confirm adoption, then switch to the annual rate once it proves out. **Pay for Microsoft 365 Copilot Enterprise** ($30 per user, plus E3 or E5) when a larger organization needs the full assistant with enterprise compliance. Budget the all-in cost and pilot before you scale. **Add Copilot Studio or Cowork** only if you are building or running agents. Treat the credits as a metered utility, set spending caps, and estimate against a real workload before you commit. It comes down to three habits: match the plan to the job, never blanket-license, and budget the base license alongside the add-on. ## Why Copilot's Price Matters for Your Brand One point outlasts any specific price. Copilot is not just a tool people work inside. When someone asks the free Copilot or Copilot Chat a question, it searches the web through Bing's index, reads pages, and answers by **citing them with clickable footnotes**. Those citations are real traffic and real authority, and they point to specific pages. So the more useful question than which plan to buy is whether Copilot names your brand and cites your page when a customer asks it what to buy. That visibility does not come on a pricing tier. In our experience auditing brands across AI engines, most companies have no idea whether Copilot can even see them, let alone whether it cites them. The full method is in our guide to [showing up in Microsoft Copilot](https://geotoolbox.ai/blog/copilot-seo). Copilot's prices will move again, so confirm the current figure on Microsoft's pricing page before you pay. And once you know what Copilot costs, the next question is what it tells people about you. Our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) maps where Copilot and the other AI engines cite sources your brand is missing from, so you can see which conversations to join, and our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) covers the broader approach. ## Frequently Asked Questions ### How much does Microsoft Copilot cost per month? Microsoft Copilot is free to start. Paid consumer plans run from $9.99 a month, with Microsoft 365 Premium at $19.99 putting Copilot in your Office apps. For work, Microsoft 365 Copilot Business is $18 to $21 per user per month and Microsoft 365 Copilot Enterprise is $30 per user per month, both on top of a required Microsoft 365 base license, as of June 2026. ### Is Microsoft Copilot free? Yes. There is a free consumer Copilot for web chat, image generation, and voice, and a free Microsoft 365 Copilot Chat for work that comes with an eligible Microsoft 365 subscription. Both answer from the public web and files you upload, but neither reads your own emails, files, or meetings on its own. That grounding only comes with the paid Microsoft 365 Copilot license. ### What's the difference between the $18 and the $30 Copilot? The $18 to $21 plan is Microsoft 365 Copilot Business, for organizations up to 300 users. The $30 plan is Microsoft 365 Copilot, the enterprise tier with no seat cap, layered on E3 or E5 licenses. Both are add-ons that require a qualifying Microsoft 365 base plan underneath. ### Do I need a Microsoft 365 subscription to use Copilot? For the free consumer Copilot, no, just a Microsoft account. For the paid business plans, yes. Microsoft 365 Copilot is an add-on that requires a qualifying Microsoft 365 base license first, so the real cost is the base plan plus the Copilot add-on, from about $24 to $90 per user per month all-in (a bundled small-business plan is the cheapest path). ### Is Copilot Pro discontinued? For new customers, yes. Microsoft stopped selling the standalone Copilot Pro subscription on October 1, 2025 and folded its features into Microsoft 365 Premium at $19.99 a month; existing Copilot Pro subscriptions ran until support ended on August 1, 2026. Guides and AI answers that still quote "Copilot Pro at $20" as the main consumer plan are out of date. ### Is Microsoft Copilot worth it compared to ChatGPT Plus? For a standalone chatbot, [ChatGPT Plus](https://geotoolbox.ai/blog/chatgpt-pricing) at around $20 is cheaper and many find it sharper. You pay the Copilot premium for something different: an assistant grounded in your own company's email, files, and meetings that works inside Word, Excel, and Outlook. If you only want a general chatbot, a $20 rival is the simpler buy. ## Sources - Microsoft 365 Copilot plans and pricing (business and enterprise) - Microsoft - `microsoft.com/en-us/microsoft-365-copilot/pricing` - Copilot pricing plans for individuals - Microsoft - `microsoft.com/en-us/microsoft-365-copilot/pricing/individuals` - Meet Microsoft 365 Premium (Copilot Pro retirement) - Microsoft - `microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse` - About Microsoft Copilot Pro (existing-subscriber wind-down) - Microsoft Support - `support.microsoft.com/en-us/microsoft-365-copilot/about-microsoft-copilot-pro` - Copilot Cowork is now generally available - Microsoft - `microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available` - Microsoft 365 Copilot pricing (Copilot Studio) - Microsoft - `microsoft.com/en-us/microsoft-365-copilot/pricing/copilot-studio` - Microsoft 365 enterprise plans and pricing (E3, E5, E7) - Microsoft - `microsoft.com/en-us/microsoft-365/enterprise/microsoft365-plans-and-pricing` - Advancing Microsoft 365: new capabilities and pricing update (July 1, 2026) - Microsoft - `microsoft.com/en-us/microsoft-365/blog/2025/12/04/advancing-microsoft-365-new-capabilities-and-pricing-update` - Microsoft 365 Copilot Pricing Calculator 2026 - Velosio - `velosio.com/blog/m365-copilot-pricing-calculator` - Claude plans and pricing - Anthropic - `claude.com/pricing` - Copilot Cowork pricing and usage details - IT Pro - `itpro.com/technology/artificial-intelligence/copilot-cowork-is-now-generally-available-everything-you-need-to-know-including-pricing-usage-limits-and-new-features` - Google AI Pro and Ultra subscriptions - Gemini - `gemini.google/subscriptions` --- ## Claude Code Context Window: What It Is and Why It Controls Costs > What the Claude Code context window is, how big it gets (200K vs 1M), what fills it, and how to read the /context meter before it degrades your output. - Canonical: https://geotoolbox.ai/blog/claude-code-context-window - Published: 2026-07-04 · Updated: 2026-07-25 The Claude Code context window is the single number that decides how much your session can hold, how much it costs, and how sharp it stays. Most people meet it only when [Claude Code](https://geotoolbox.ai/blog/what-is-claude-code) starts forgetting things or a usage limit lands mid-week. This is what the context window actually is, how big it is on each model, what fills it before you type a word, and how to read and keep it under control. ## What the Claude Code Context Window Is The context window is the session's working memory: everything Claude Code can see at one moment, measured in tokens. That includes the system prompt, your project's `CLAUDE.md`, the tools and MCP servers you have connected, every file it has read, every tool result, and the full back-and-forth of your conversation. It all shares one fixed budget, and it all competes for the same space. A [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai) is roughly four characters, or about three-quarters of a word. Code is denser than prose because of punctuation, identifiers, and syntax, so source files eat the window faster than their line count suggests. Reading one large module can cost tens of thousands of tokens on its own. The mental model that trips people up is treating the window like human memory, something Claude holds in the back of its mind and glances at when needed. It does not work that way. The model has no memory between turns, so Claude Code re-sends the entire window on every single request, then appends your new message. It is not a filing cabinet Claude opens occasionally; it is the whole document it re-reads, start to finish, every time you hit enter. That is what makes everything below matter: the [context window](https://geotoolbox.ai/glossary/context-window) is not just a capacity limit, it is a recurring cost you pay on every turn. ## How Big Is the Context Window? 200K vs 1M **1 million tokens is now the normal case, not the exception.** Every current model except Haiku 4.5 is 1M-capable, and Sonnet 5 runs at 1M unconditionally. The 200,000-token window is still the documented default, but in practice it is the floor you land on in specific situations rather than the standard you start from.
ModelContext window in Claude CodeNotes
Claude Sonnet 5 (current Sonnet)1M tokensAlways 1M on the Anthropic API. No 200K variant, no [1m] suffix, no usage credits on any paid plan
Claude Opus 5 (current Opus)1M tokensAutomatic on Max, Team, and Enterprise. Pro must enable usage credits
Fable 51M tokensMax/Team Premium (Pro via usage credits)
Haiku 4.5200K tokensThe one current model that does not reach 1M
Legacy (Opus 4.8 / 4.7 / 4.6, Sonnet 4.6)1M tokensOpus 4.8 and 4.7 follow the Opus rules above. Sonnet 4.6 is not part of the automatic upgrade and requires usage credits on every subscription plan, including Max
The 1M window is available on Pro, Max, Team, and Enterprise, not the free tier, per Anthropic's [context window documentation](https://support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans). The plan split, from Anthropic's [model configuration docs](https://code.claude.com/docs/en/model-config): on Max, Team, and Enterprise, Opus is automatically upgraded to 1M with no configuration at all, including on both Team Standard and Team Premium seats. On Pro, Opus needs usage credits enabled first. If you run [Sonnet 5](https://geotoolbox.ai/blog/claude-sonnet-5), none of that applies, you simply get the million-token window. Two details that matter for cost. First, **the 1M window uses standard model pricing with no premium for tokens beyond 200K.** Where extended context is included with your subscription, it stays covered by your subscription; where it comes through usage credits, those tokens bill to credits. A bigger window costs more because you are sending more tokens, not because the tokens themselves are priced higher. Second, you can select it explicitly with the `[1m]` suffix: ``` /model opus[1m] /model claude-opus-5[1m] ``` If your account supports the 1M window, it also appears in the `/model` picker; restart the session if you do not see it. `sonnet[1m]` has no effect when `sonnet` already resolves to Sonnet 5, since that model is 1M natively. You land back on 200K in three cases: running Haiku 4.5, running behind an LLM gateway (where Claude Code cannot verify 1M support), or setting `CLAUDE_CODE_DISABLE_1M_CONTEXT=1`. Two things about that headline number. First, these are the Claude Code sizes; the Claude.ai chat app is different. The [same documentation](https://support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans) now gives both Opus 5 and Sonnet 5 a 1M window in chat on any paid plan, while older model versions are held to 500K there, so do not carry a chat figure over to your terminal without checking which model it refers to. Second, the number you can actually use is always smaller than the ceiling. By the time a real session is running, a chunk of the window is already spoken for by things you never typed, which is the next section. ## What Fills the Window Before You Type a Word Open a fresh Claude Code session and the window is already partly full. A stack of components loads automatically, before your first message. The figures below are the representative values Anthropic publishes in its own [context window simulator](https://code.claude.com/docs/en/context-window); your setup will vary, particularly on `CLAUDE.md` size and skill count.
What loads automaticallyRepresentative tokens
System prompt (Claude Code's own instructions and tool definitions)4,200
Project CLAUDE.md1,800
Auto-memory (MEMORY.md from past sessions)680
Skill descriptions450
Personal ~/.claude/CLAUDE.md320
Environment info280
MCP tool names (schemas deferred until used)120
Two notes on that table. Anthropic folds tool definitions into the system prompt rather than counting them separately, so a single 4,200-token figure covers both. And auto-memory is capped at **the first 200 lines or 25KB of `MEMORY.md`, whichever comes first**, which in practice lands well under a thousand tokens rather than the several thousand the cap might suggest. Git branch, status, and recent commits load as a separate block at the very end of the system prompt. None of that is waste, but it is fixed overhead on every turn, and two sources balloon fast if you are not watching. ### MCP Servers, the Big One The first is MCP servers. Each connected server advertises its tools, and a single tool definition, name plus description plus parameter schema, can run several hundred tokens. Back in December 2025, before deferral existed, developer Damian Galarza [measured one MCP tool at roughly 663 tokens](https://www.damiangalarza.com/posts/2025-12-08-understanding-claude-code-context-window/) and the Playwright MCP's 22 tool definitions at about 14,300 tokens, over 7 percent of a 200K window. That is the problem Anthropic solved: **MCP tool definitions are now deferred by default** and loaded on demand through tool search, so only tool names consume context until Claude actually reaches for a specific tool. Run `/mcp` to see per-server costs. One documented exception. When `ANTHROPIC_BASE_URL` points at a non-first-party host, MCP tool search is disabled by default; set `ENABLE_TOOL_SEARCH=true` if your proxy forwards `tool_reference` blocks. Outside that case, deferral is simply on. Names still add up across many servers, though, so servers you never touch remain weight you are not using. ### Skills Load Lighter Than MCP The second is skills. A [skill](https://geotoolbox.ai/blog/how-does-claude-work) only loads its name and description at startup, a couple hundred tokens, and pulls in the full instructions on demand when invoked. That is the same deferral idea applied more aggressively: the Playwright capability that costs 14,300 tokens as a full MCP definition is about 200 tokens as a skill until you actually use it. If a tool is one you reach for occasionally rather than every session, a skill is far cheaper on context. You can go further and keep even the description out: `disable-model-invocation: true` in a skill's frontmatter means nothing loads until you call it explicitly with `/name`, and `skillOverrides` in settings does the same for skills you did not write. This is why "audit your MCP servers" is the single biggest context cleanup most people have never done. If you are deciding which servers are worth their weight in the first place, our roundup of [open-source tools for Claude Code](https://geotoolbox.ai/blog/claude-code-open-source-tools) sorts them by the job they do. ## How to Read the /context Meter You do not have to guess at any of this. Run `/context` in a session and Claude Code prints a live map of the window: a filled grid plus a per-category breakdown of exactly what is taking up space.
![An annotated Claude Code /context readout showing the window split into system prompt, system tools, MCP tools, memory files, messages, free space, and the reserved autocompact buffer, each with its token count and percentage.](/blog/claude-code-context-window/context-meter-breakdown.png)
The /context readout breaks the window into categories. The reserved autocompact buffer and free space are the two numbers to watch.
The top line shows your model and usage, something like `101k/200k tokens (51%)` — on a current model the denominator will read against the 1M window instead. Below it, each category reports its own tokens and percentage: system prompt, system tools, MCP tools, custom agents, memory files, messages, then free space and a reserved autocompact buffer at the bottom. The value is that it turns a vague "why is this session sluggish" into a specific answer. If MCP tools are showing 26k tokens, you know where to cut. If your `CLAUDE.md` is 4k before you have done anything, that is a fixed tax on every turn worth trimming. ### The Interactive Simulator Anthropic recently shipped an [interactive simulator](https://code.claude.com/docs/en/context-window) that walks through the same thing visually: what loads at startup, what each file read costs, and when rules and hooks fire as a session grows. It is the fastest way to build an intuition for how quickly the window fills before you spend real tokens learning it the hard way. One myth to put down: older guides claim Claude Code has no native way to see token usage and that you only notice trouble through behavior. That is out of date. Between `/context`, the `/usage` command (which on paid plans breaks recent usage down by skills, subagents, plugins, and individual MCP servers), and a [configurable status line](https://code.claude.com/docs/en/costs) that shows context usage continuously, the window is fully visible. You just have to look. ## Auto-Compaction: What Happens When the Window Fills Claude Code does not crash when you approach the limit. It compacts, and it does so in two stages: **it clears older tool outputs first, then summarizes the conversation if that is not enough.** Your requests and key code snippets are preserved; detailed instructions from early in the conversation may not be. That reserved band at the bottom of the `/context` readout, the autocompact buffer, is the room held back so there is space to run the summarization when it triggers. Anthropic does not publish a fixed buffer size, but the trigger is now configurable. `CLAUDE_CODE_AUTO_COMPACT_WINDOW` sets the capacity used for auto-compaction calculations, defaulting to the model's context window, except on Sonnet 5, which auto-compacts at about **967K tokens** by default. `CLAUDE_AUTOCOMPACT_PCT_OVERRIDE` sets the percentage of that window at which compaction fires. On a local Opus session, auto-compaction triggers when the conversation reaches the model's context limit rather than at a percentage. There is one failure mode worth recognizing. If a single file or tool output is so large that the context refills immediately after each summary, Claude Code **stops auto-compacting after a few attempts and shows an error instead of looping**. If you see that, the fix is a smaller input, not a bigger window. ### What Survives a Compaction Compaction is not lossless, and it is worth knowing what survives it. Anthropic's [documentation on what survives compaction](https://code.claude.com/docs/en/context-window) notes that the system prompt stays unchanged and isn't part of the compacted message history in the first place, while your project-root `CLAUDE.md`, memory, and MCP tools all reload automatically. **The skill listing is the one exception.** Everything else reloads; the index of available skill descriptions does not. Only skills you actually invoked are preserved, and those are capped at about 5,000 tokens each and 25,000 total, oldest dropped first. Truncation keeps the start of the file, which is why the most important instructions belong near the top of a `SKILL.md`. Path-scoped rules and nested `CLAUDE.md` files are dropped until their directory is touched again. In other words, the standing context you set up mostly holds, but anything that depended on being deep in the conversation can quietly vanish. ### /clear, /compact, and /rewind You have three manual controls, and they are not interchangeable.
CommandWhat it doesReach for it when
/clearWipes the conversation entirely, keeps your files and toolsYou are switching to unrelated work
/compactSummarizes the conversation and continues from the summaryYou hit a natural break in the same task
/rewindRestores the conversation or code to an earlier checkpointA path went wrong and you want to back out of it
The trick with `/compact` is to run it before quality drops, not after. Compact at a clean boundary, say once a phase of work is done, and the summary is built from good state. Compact a session that has already gone off the rails and you bake the confusion into the summary. You can also steer it: `/compact Focus on the auth refactor and the failing tests` tells Claude what to keep. If you find yourself typing the same steer every time, make it permanent with a `# Compact instructions` section in your project-root `CLAUDE.md`, and every compaction in that project inherits it. Reach for `/clear` between tasks and `/compact` within one. ## Why a Full Window Degrades Your Output A window that is technically not full can still produce worse output. This is where the context window quietly starts costing you quality on top of tokens. The reason is how attention works. Models weight some positions more heavily than others, and content in the middle of a long window gets less reliable attention than content at the start or end, the well-documented "lost in the middle" effect. So the coding conventions you set an hour ago drift toward the middle as the session grows and start getting ignored. The symptoms are recognizable: Claude repeats work it already did, renames a function it just settled on, reintroduces a pattern you told it to avoid, or asks a question you already answered. None of that is a bug. It is a noisy window crowding out the signal. The trap is assuming this only kicks in when the meter is near full. It does not. Degradation is gradual, and on a busy session it can show up while you still have plenty of free space, because what matters is not the percentage used but how much noise sits between the model and the thing it needs to focus on. Treat the `/context` percentage as a rough gauge, not a green light to keep piling on until it fills. ### Why a Bigger Window Isn't the Fix This is also why a 1M window is not the fix it sounds like. Chroma's [Context Rot study](https://www.trychroma.com/research/context-rot) tested 18 frontier models, Claude among them, and found every one degrades as input length grows, even on simple retrieval tasks. A bigger window buys you more room before you have to compact. It does not buy immunity from the model getting less reliable the more you cram in. Curating what goes into the window beats maximizing how much fits. The cost angle compounds here. A bloated session is expensive twice: you pay to re-send that growing window on every turn, and you pay again when auto-compaction spends tokens summarizing it, sometimes summarizing a previous summary. In our experience, the sessions that burn the most tokens are rarely the hardest problems. They are ordinary tasks run in a window nobody cleared, where each turn re-reads a pile of stale context that stopped being useful long ago. Computerphile has a good walkthrough of why re-processing that context is expensive at the hardware level: Video: https://www.youtube.com/watch?v=-0HRzXk8vlk ## Context Window vs Usage Limit These two get conflated constantly, and the difference matters. The **context window** is per-session working memory, the 200K or 1M budget for one conversation, and `/clear` resets it to empty. Your **usage limit** is a plan-level cap on how much you can consume over time, a weekly or monthly ceiling on Pro or Max. Clearing your context does nothing to your usage limit directly; it just makes each turn cheaper, which means you reach that limit slower. So "saving tokens" means two different things depending on which one is biting. If you are hitting a usage limit mid-week, the goal is fewer tokens consumed overall. If a single session is degrading, the goal is a leaner window. The habits overlap, but the reason you care is different. For the full set of levers on the spend-and-limit side, model routing, reasoning effort, and prompt discipline, see our companion guide on [reducing Claude Code token costs](https://geotoolbox.ai/blog/reduce-claude-code-token-costs). For how the plans and usage credits themselves work, our [Claude pricing breakdown](https://geotoolbox.ai/blog/claude-pricing) covers the tiers. ## Keeping the Window Lean Managing the window well is mostly a handful of habits, not a tool. Clearing and compacting deliberately, covered above, is half of it. The other half is controlling what enters the window in the first place. Point Claude at specific files with `@path` or line ranges instead of asking it to "figure out the codebase," one of the fastest ways to drain the window. Hand noisy jobs like log analysis or codebase exploration to a subagent, which works in its own separate window and returns only a summary. Keep `CLAUDE.md` tight, Anthropic suggests under 200 lines, and move detailed per-workflow instructions into skills so they load on demand instead of on every turn. And run `/mcp` to switch off servers you are not using. Where a CLI and an MCP server do the same job, the CLI (`gh`, `aws`, `gcloud`) is cheaper on context because it adds no per-tool listing at all. ### Agent Teams Multiply the Bill One feature deserves its own warning if you are watching spend. Agent teams, disabled by default and enabled with `CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1`, use **roughly 7x more tokens** than a standard session when teammates run in plan mode. Each teammate is a separate Claude instance maintaining its own context window, and each one loads `CLAUDE.md`, MCP servers, and skills automatically. Anthropic's own mitigations: run teammates on Sonnet, keep teams small, keep spawn prompts focused, and shut teammates down when they are done. It is a real capability, but it is the most expensive button in the product. One model-specific wrinkle: extended thinking is billed as output, so those [reasoning](https://geotoolbox.ai/glossary/reasoning-model) tokens count against your budget too. On [Fable 5](https://geotoolbox.ai/blog/fable-5-ban) you cannot turn extended thinking off, only lower the effort, so it takes a bigger bite of the window on simple tasks. ## The Window Is the Lever The context window is not a background detail. It is the resource that decides both what a session costs and how good its output stays. Watch it with `/context`, keep it lean, and compact deliberately, and most of the "why is Claude Code expensive and getting worse" problem takes care of itself. The reason a coding agent's context is expensive is the same reason serving AI answers is expensive: models re-process large amounts of context to produce each response, and AI search engines do much the same with your pages when they decide whether to cite you. That is the side we work on at geotoolbox. If you run a site and want to see how AI models read and reference it, our free [AI readiness checker](https://geotoolbox.ai/tools/ai-readiness) is a good place to start. ## Frequently Asked Questions ### What is the Claude Code context window size? On current models it is 1 million tokens. Sonnet 5 runs at 1M unconditionally, Opus 5 and Fable 5 reach 1M on paid plans, and Haiku 4.5 is the one current model still capped at the 200,000-token standard window. The usable space is always smaller than the headline number, because the system prompt, tools, memory, and your `CLAUDE.md` occupy part of the window before you type anything. ### Does Claude Code have a 1 million token context window? Yes. Sonnet 5 has it on the Anthropic API with no configuration and no usage credits on any paid plan. Opus is automatically upgraded to 1M on Max, Team, and Enterprise; on Pro it requires enabling usage credits first. You can select it explicitly with `/model opus[1m]`. It is not available on the free tier, and the 1M window carries no price premium for tokens beyond 200K. ### What is the difference between /compact and /clear? `/clear` wipes the conversation history entirely while keeping your files and tools, which is what you want when switching to unrelated work. `/compact` summarizes the conversation so far and continues from that summary, which is what you want at a natural break in the same task. The short version: between tasks, clear; inside one task, compact. ### Why does Claude Code get worse in long sessions? As the window fills, model attention weights recent content more heavily and instructions from earlier drift toward the middle, where they get less reliable attention. You see repeated work, contradicted decisions, and ignored rules. Research on context rot shows this happens even with a larger window, so a 1M context delays the problem rather than removing it. ### Is the context window the same as my usage limit? No. The context window is the per-session working-memory budget, and `/clear` resets it. Your usage limit is a plan-level cap on how much you can consume over a week or month. Keeping the window lean makes each turn cheaper, which helps you reach the usage limit slower, but the two are separate things. ### How do I check how much context I am using? Run `/context` in a session for a live breakdown by category, use `/usage` for a plan-limit view of recent usage by skills, subagents, plugins, and MCP servers, or configure the status line to display context usage continuously. Claude Code shows all of this natively. ## Sources - Explore the context window (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/context-window` - Model configuration (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/model-config` - How Claude Code works (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/how-claude-code-works` - Environment variables (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/env-vars` - Context windows (Anthropic, Claude Platform Docs) - `platform.claude.com/docs/en/build-with-claude/context-windows` - Manage costs effectively (Anthropic, Claude Code Docs) - `code.claude.com/docs/en/costs` - How large is the context window on paid Claude plans? (Anthropic) - `support.claude.com/en/articles/8606394-how-large-is-the-context-window-on-paid-claude-plans` - Understanding the Claude Code context window (Damian Galarza) - `damiangalarza.com/posts/2025-12-08-understanding-claude-code-context-window` - Context Rot: How Increasing Input Tokens Impacts LLM Performance (Chroma Research) - `trychroma.com/research/context-rot` - Why AI Tokens are so Expensive (Computerphile) - `youtube.com/watch?v=-0HRzXk8vlk` --- ## How to Reduce Claude Code Token Costs (and Why They Add Up Fast) > Claude Code burns tokens fast because it re-reads your whole context on every step. Here is the mechanism, plus the highest-impact ways to cut your token bill. - Canonical: https://geotoolbox.ai/blog/reduce-claude-code-token-costs - Published: 2026-07-04 · Updated: 2026-08-14 You open [Claude Code](https://geotoolbox.ai/blog/what-is-claude-code) on a Monday, ask it to fix what looks like a two-file bug, and by lunch you have hit your weekly usage limit. You are not doing anything wrong, and you are not alone. Cost is one of the most common complaints about agentic coding tools right now, sharpened by [Fable 5 coming back online](https://geotoolbox.ai/blog/fable-5-ban) in July 2026 and a lot of people rediscovering how fast a coding agent can spend. Token cost is not random. It mostly comes down to one mechanism, and once you understand it, a handful of habits cut the bill hard. This guide covers why Claude Code burns tokens so fast, then the levers that actually move the number, ranked by how much they save. ## Why Claude Code Burns Tokens So Fast The model has no memory between steps. Every time Claude Code sends a request, it re-sends the entire context: the system prompt, your project files, every prior message, and every tool result, with your new message appended at the end. Anthropic's own [prompt caching documentation](https://code.claude.com/docs/en/prompt-caching) states it plainly: "The model doesn't remember anything between requests, so Claude Code re-sends the full context." A human holds a conversation in their head. The model re-reads the whole transcript, out loud, on every single turn. That is expensive on its own. Agentic coding makes it worse. When you ask for that "simple" fix, Claude does not answer once. It fetches a diff, reads six files to understand them, runs the tests, reads the failures, and edits. Each of those steps injects its full output, entire files, thousands of lines of logs, back into the context. And the whole growing pile is re-sent on the next step, and the next. This is why input tokens, not the code Claude writes, dominate your cost by volume. It is also the same [transformer](https://geotoolbox.ai/glossary/transformer-model) machinery that powers every model, so the pattern is worth understanding once: see [how Claude actually works](https://geotoolbox.ai/blog/how-does-claude-work) for the longer version, and [what tokens are](https://geotoolbox.ai/blog/what-are-tokens-in-ai) if that word is still fuzzy. Claude Code softens this with [automatic prompt caching](https://code.claude.com/docs/en/prompt-caching): the unchanged part of your context is re-billed at roughly a tenth of the normal input rate, so you are not paying full price to re-read a stable prefix every turn. That shifts the real cost onto what changes each turn, plus anything that throws the cache away. Keeping that cache intact is a lever in itself, one we come back to below.
![Across four turns, the re-sent conversation history keeps accumulating while only a small new slice is added each turn, so the re-sent part dominates the token bill.](/blog/reduce-claude-code-token-costs/context-grows-every-turn.png)
Every step re-sends the entire history plus a small new slice, so the re-sent part dominates the tokens you send and grows each turn.
The scale is not subtle. A 2026 study, [How Do AI Agents Spend Your Money?](https://arxiv.org/abs/2604.22750), measured agentic coding tasks consuming roughly 1000 times more tokens than ordinary code chat, with input tokens driving the cost. The same study found the exact same task can vary by up to 30 times in token usage from one run to the next. A task that looks trivial to you can quietly turn into fifty thousand tokens of re-read context. That variance is why two developers on the same team can see wildly different bills, and why "it was just a small change" is never a reliable predictor of cost. Computerphile has a clear walkthrough of why this is so expensive at the hardware level: Video: https://www.youtube.com/watch?v=-0HRzXk8vlk ## The Sonnet 5 Tokenizer Tax One recent change can raise your bill without you touching anything. [Sonnet 5](https://geotoolbox.ai/blog/claude-sonnet-5) launched on June 30, 2026 and is now the default model in Claude Code, and it ships with a new tokenizer that emits roughly 30 percent more tokens for the same text on average, per Anthropic's [what's new in Sonnet 5](https://platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5) doc. The inflation is uneven: heavier on English prose, on the order of 40 percent, and lighter on code, closer to 27 percent on Python, with languages like Chinese roughly unchanged. Per-token pricing did not change, so an equivalent request or session simply costs about a third more tokens than the same work did on Sonnet 4.6. The coverage calls it a "tokenizer tax," and the name fits. You cannot opt out of the ratio, but it raises the stakes on every lever below. If the same task now costs roughly 30 percent more tokens, the context you keep lean and the output you keep short are each worth proportionally more, and the 1M-token window holds meaningfully less, roughly a quarter less, of your actual text than it used to, so context discipline starts biting sooner. ## Your Bill or Your Limit? What "Cheaper" Actually Means Before the tactics, get clear on what you are optimizing, because it changes what "saving tokens" means for you. If you use Claude Code through a Pro or Max subscription, you do not pay per token. You pay a flat monthly fee and hit a usage limit. Saving tokens means not running out mid-week, not a smaller invoice. If you use an API key, every token is real money, and the numbers add up: Anthropic's [cost documentation](https://code.claude.com/docs/en/costs) reports an average of around 13 US dollars per developer per active day and 150 to 250 dollars per developer per month across enterprise deployments, with 90 percent of users staying under 30 dollars a day. Either way, the same habits help. One protects your limit, the other protects your wallet. You can see where your tokens go with the `/usage` command. It shows your session token totals, and on a paid plan (Pro, Max, Team, or Enterprise) it also breaks recent usage down by skills, subagents, plugins, and individual MCP servers: ``` Total cost: $0.55 Total duration (API): 6m 19.7s Total duration (wall): 6h 33m 10.2s ``` The dollar figure is a local estimate, not your actual bill, but the breakdown is what matters: it tells you which parts of your setup are quietly eating context. For the full plan comparison, see our guide to [Claude pricing in 2026](https://geotoolbox.ai/blog/claude-pricing). Look at that breakdown once before you start cutting, so you are trimming real waste rather than guessing. ## The Levers That Actually Move Your Token Bill Most guides hand you 30 or 50 tips in no particular order. That is not useful when you have five minutes and a limit warning. Almost all of the savings come from two things: keeping your context small and matching the model to the task. Everything else is a smaller, situational win. Here is the ranking, biggest wins first.
LeverWhat it doesEffortTypical impact
Manage context (/clear, /compact)Stops re-sending stale history on every turnLowHighest
Match the model to the taskRoutine work on Sonnet or Haiku, not OpusLowHigh
Lower reasoning effortCuts thinking tokens on simple tasksLowHigh
Disable unused MCP serversRemoves tool definitions from contextLowMedium
Filter verbose command outputKeeps logs and diffs out of contextMediumMedium
Write specific promptsAvoids broad scanning and re-workLowMedium
Delegate noisy work to subagentsIsolates verbose output from the main threadMediumSituational
Start at the top. The next sections take the levers in that order. ## Keep Your Context Small This is the lever that matters most, because it attacks the mechanism directly. Every token you keep out of the [context window](https://geotoolbox.ai/glossary/context-window) is a token you do not pay to re-read on every future turn. For the mechanics of the window itself, how big it is and how to read the `/context` meter, see our guide to the [Claude Code context window](https://geotoolbox.ai/blog/claude-code-context-window). Two commands do most of the work, and people confuse them. `/clear` wipes the conversation history entirely while keeping your files and tools. Use it when you switch to unrelated work, so the React bug you just fixed stops riding along into a Postgres question. `/compact` replaces a long back-and-forth with a short summary and keeps going, preserving the decisions and current task state while discarding raw tool output. Use it at a natural break, when you finish a phase, not reactively after Claude starts forgetting things. A healthy session compacts into a better summary than a degraded one. The bigger mistake is letting a session ride all the way to the context ceiling. At that point Claude Code auto-compacts for you, spending tokens to summarize and then carrying a lossy version forward, and a session you keep nursing past that point pays for that again and again while getting fuzzier each time. When you do not actually need the history, a clean `/clear` or a fresh session is cheaper and sharper than riding repeated auto-compactions. You can tell compaction what to keep: ``` /compact Focus on the auth refactor and the failing tests, drop the exploration ``` Run `/context` any time to see what is actually occupying the window: open files, tool definitions, conversation turns. It is a memory profiler for your session, and it is the fastest way to spot a file that got pulled in and never left. Two habits keep the window lean from the start. First, do not let Claude scan the whole repo. Point it at the specific files with `@path/to/file` instead of asking it to "figure out the codebase," which is one of the biggest single token drains there is. Second, keep your `CLAUDE.md` small. It loads on every session and sits in context the whole time, so a bloated one is a fixed tax on every turn. Anthropic suggests keeping it under 200 lines and moving specialized workflows into skills, which load only when invoked: ```markdown # Compact instructions When compacting, focus on test output and code changes. ``` If you have gone down a wrong path, `/rewind` back to an earlier checkpoint instead of arguing Claude out of it. Rewinding truncates to context that is already cached, so it is usually cheaper than compacting a session full of dead ends. ## Match the Model to the Task Most people open Claude Code, leave it on the default model, and never think about it again. Running everything on Opus is where a lot of bills quietly balloon. Opus is the architect: reserve it for planning, hard debugging, and decisions where the reasoning is the point. [Sonnet 5](https://geotoolbox.ai/blog/claude-sonnet-5) is the contractor that handles day-to-day implementation, refactors, and tests. Haiku is for cheap helper work and exploration. Use `/model` to switch, or set a default in `/config`.
ModelBest forRelative cost
OpusArchitecture, planning, hard reasoningHighest
SonnetEveryday coding, refactors, tests, reviewsMiddle
HaikuSimple edits, questions, subagent helpersLowest
The "opusplan" setting automates the split: it plans on Opus, then drops to Sonnet for execution. Useful, with one caveat worth knowing. Each model and each effort level has its own cache, so switching mid-session recomputes the entire request with no cache hits. That is why picking your model and effort at the top of a session beats flipping between them mid-task. The other model-shaped lever is reasoning effort. Extended thinking is on by default, and those thinking tokens are billed as output, with a default budget that can run into tens of thousands of tokens. For a job like "rename this variable," that is pure waste. Drop it with `/effort` for simple tasks, or, on models with a fixed thinking budget, cap it with `MAX_THINKING_TOKENS`. Heavier [reasoning](https://geotoolbox.ai/glossary/reasoning-model) is a real trade-off, not a free upgrade, so dial it down when the task does not need it. One note for the current lineup: on Fable 5 you cannot turn extended thinking off entirely, so that particular switch is off the table, though you can still lower the effort. ## Cut the Hidden Context Bloat Some of the worst offenders never show up in anything you typed. They load quietly in the background. MCP servers are the classic example. Older advice says every connected server dumps its full tool definitions into context on every message, and for a while that was true. It is now more nuanced: on supported models, tool definitions are deferred by default, so only tool names enter context until Claude actually uses one. That said, disabling servers you are not using still helps, especially on Haiku or through some gateways where deferral is unavailable and the full definitions load. Run `/mcp` to see what is connected and turn off what you do not need. When a plain command line tool like `gh` or `aws` will do, prefer it, since it adds no per-tool listing at all. For the servers that do earn their place, plus the cost trackers that keep this spend visible, see our roundup of [open-source Claude Code tools](https://geotoolbox.ai/blog/claude-code-open-source-tools). Verbose command output is the other silent drain. A test run, a `git diff`, or a build log gets billed at full token cost, and most of it is noise Claude does not need. Filter it before it lands. That can be as simple as `pytest -q`, or a preprocessing hook that greps a 10,000-line log down to the errors, turning tens of thousands of tokens into hundreds. The output nobody thought to trim is almost always where the surprise token counts come from, not the prompts people worry about. For genuinely noisy jobs, hand them to a subagent. A subagent explores logs, runs the tests, or reads the docs in its own separate context, and only a short summary comes back to your main thread. The verbose part stays quarantined, and you can pin the subagent to a cheaper model like Haiku for that grunt work. Finally, use plan mode (Shift+Tab) before a non-trivial task. Claude proposes an approach before touching anything, which heads off the expensive re-work of letting it charge down the wrong path first. ## Write Prompts That Don't Waste Tokens How you phrase a request changes how many tokens come back. Vague prompts like "improve this codebase" trigger broad scanning and long, hedged answers. Specific ones keep the work tight: "add input validation to the login function in auth.ts" lets Claude read one file and move. Negative prompts help too. Adding "do not redesign the architecture, do not add dependencies, do not touch code outside this function" reduces collateral changes and the tokens spent making them. Answer budgets are the cheapest lever of all, and they target the priciest tokens. Output tokens are billed five times more than input tokens, per token: [Anthropic's pricing](https://platform.claude.com/docs/en/about-claude/pricing) puts Sonnet 5 at 2 dollars per million input against 10 dollars per million output per million output, a 5-to-1 ratio. So "return only the diff," "answer in under 200 tokens," or "five bullets maximum" saves more than the word count suggests, and it usually improves the answer by forcing Claude to select instead of generating everything plausible. We lean on this constantly in our own work. On routine tasks we ask for terse output by default and let the model spend its budget on reasoning rather than narration. The point is to constrain the output, not the thinking: a short answer is not a dumber one if the model still reasoned fully before replying. One caveat, since extended thinking is billed as output too: an answer-length limit does not cap the hidden thinking tokens, so on simple tasks pair the terse-output request with a lower `/effort`. Some developers push this further into a grammar-stripped "caveman" style for quick lookups, which saves input tokens too, with one real failure mode. Strip too much and Claude infers more, which means it sometimes infers wrong. Keep full sentences for anything that touches production code. One developer documented a roughly 60 percent cut from this kind of context and prompt discipline alone, in [this dev.to write-up](https://dev.to/numbpill3d/how-i-cut-my-claude-code-token-usage-by-60-and-got-better-output-48b0). For sequences you run constantly, define a `/command` so a repeated workflow stays one short, consistent instruction instead of a long ad-hoc prompt retyped each time. It still costs tokens when invoked, but a tight command beats a sprawling one. ## The One Habit Worth Keeping The habit that pays back on every turn is context discipline. Clear between tasks, compact at the breaks, keep `CLAUDE.md` lean, and stop the agent from re-reading things it does not need. Model routing and sharper prompts stack on top, but managing context is the lever that keeps working while you do everything else. The reason agentic coding is expensive is the same reason serving AI answers is expensive: models re-process large amounts of context to produce each response. We spend our time on the visibility side of that economics at geotoolbox, so if you also run a site and want to see how AI models read and cite it, our free [AI readiness checker](https://geotoolbox.ai/tools/ai-readiness) is a good place to start. ## Frequently Asked Questions ### Why does Claude Code use so many tokens? Because the model has no memory between steps, so Claude Code re-sends the entire context, your files, prior messages, and every tool result, on every single turn. Agentic tasks compound this by injecting large file and log outputs that then ride along on all following turns, which is why input tokens dominate the cost by volume. ### What is the difference between /clear and /compact? `/clear` wipes the conversation history completely while keeping your files and tools, which is what you want when switching to unrelated work. `/compact` summarizes the conversation so far and continues from that summary, which is what you want at a natural break in the same task. Reach for `/clear` between tasks and `/compact` within one. ### How big should my CLAUDE.md file be? Keep it small, ideally under 200 lines. It loads at session start and stays in context the entire session, so every line adds fixed weight to every task, including the ones it has nothing to do with. Move detailed or specialized workflow instructions into skills, which load only when Claude invokes them. ### Does switching from Opus to Sonnet actually lower my bill? Yes, for routine work. Opus costs more per token, so running everyday edits, refactors, and tests on Sonnet (or simple helpers on Haiku) cuts cost with no real quality loss. Reserve Opus for planning and hard reasoning. Set your model at the start of a session, since switching mid-session recomputes the cache. ### Do subagents cost more or less? It depends on the job. A subagent runs verbose work like log analysis in its own separate context and returns only a short summary, which keeps that noise out of your main thread and can save tokens overall. Spawn many long-running teammates at once, though, and each carries its own context window, so cost scales with the team. ### Is reducing token usage about my bill or my usage limit? Both, depending on your plan. On a Pro or Max subscription you pay a flat fee and hit a usage limit, so saving tokens means not running out mid-week. On an API key every token is billed, so it is literal money. The same habits help either way. ## Sources - Manage costs effectively - Claude Code Docs (Anthropic) - `code.claude.com/docs/en/costs` - How Claude Code uses prompt caching (Anthropic) - `code.claude.com/docs/en/prompt-caching` - How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks (arXiv, 2026) - `arxiv.org/abs/2604.22750` - Why AI Tokens are so Expensive (Computerphile) - `youtube.com/watch?v=-0HRzXk8vlk` - Claude API and model pricing (Anthropic) - `platform.claude.com/docs/en/about-claude/pricing` - What's new in Claude Sonnet 5 (Anthropic) - `platform.claude.com/docs/en/about-claude/models/whats-new-sonnet-5` - How I Cut My Claude Code Token Usage by 60% and Got Better Output (dev.to) - `dev.to/numbpill3d/how-i-cut-my-claude-code-token-usage-by-60-and-got-better-output-48b0` --- ## Open Weights vs Open Source: The Real Difference (2026) > Open weights and open source are not the same thing. See what each term means, which 2026 AI models are which, the license catches, and why it matters. - Canonical: https://geotoolbox.ai/blog/open-weights-vs-open-source - Published: 2026-07-03 · Updated: 2026-08-19 Open weights and open source get used as if they mean the same thing. They do not. Open weights means you can download the finished model and run it. Open source means you also get the recipe: the training code and enough detail about the data to rebuild and audit the model from scratch. Almost every model marketed as "open source" today - Llama, DeepSeek, Qwen, Gemma - is actually only open-weight. This guide covers what each term really means, which 2026 models fall where, the license catches that bite in production, and why the distinction matters if you care about being cited in AI answers. ## Open Weights vs Open Source: The Short Answer An [open-weights model](https://geotoolbox.ai/glossary/open-weights) hands you the trained model. An open-source model hands you the model plus everything needed to reproduce it. Both let you download and run the thing. The difference is what you can do beyond running it: study how it was built, verify what went into it, and rebuild it yourself if you need to. Open weights give you the product. Open source gives you the product and the factory.
What you getOpen weightOpen source (OSAID)
Trained weights to download and runYesYes
Fine-tune on your own dataYesYes
Training codeNoYes
Training data (or detailed data information)NoYes
Rebuild from scratch and fully auditNoYes
Typical licensePermissive or custom (Apache, MIT, or vendor terms)OSI-approved (e.g. Apache 2.0)
That single distinction drives everything else: what you can legally do with the model, whether you can trust what is inside it, and whether the "open" label means anything at all. And notice what the license column does not settle. The same permissive license, Apache 2.0, turns up on both sides in the model table below. What separates open weight from open source is not the license name, it is whether the training code and data come with it. ## What "Open Weights" Actually Means The weights are the numbers a model learned during training. Billions of parameters that, together, decide how it turns your prompt into a response. When a lab publishes those weights, you can download them, run the model on your own hardware, and [fine-tune](https://geotoolbox.ai/glossary/fine-tuning) it on your own data. That is a real and useful kind of openness. You are no longer renting the model through someone else's API on their meter and their terms. What you do not get is how it was made. The training code, the exact data it was trained on, the filtering and cleaning steps, the specific recipe - all of that usually stays private. You have the finished parameters, not the process that produced them. A common way to describe this: the weights are the compiled output, not the source. It is like shipping a program as an executable instead of the source code. You can run it and even patch it, but you cannot read the logic that built it or recreate it cleanly. The comparison is not perfect, and some engineers push back on it, since in practice most people work with a model by adjusting weights, not by retraining from raw data. But it captures the core limit. You are trusting the creator's choices about what went in, because you cannot see them. In practice, open weights show up on hubs like Hugging Face, where you download a model file and a license, run it locally or on rented GPUs, and adapt it. Downloading the weights is easy. Whether you can use them the way you want is a separate question, and it lives entirely in the license. ## What "Open Source AI" Actually Means Open source has a specific definition, and it is stricter than "you can download it." The Open Source Initiative, the same body that has defined open source software since 1998, published its [Open Source AI Definition](https://opensource.org/ai/open-source-ai-definition) (OSAID 1.0) on October 28, 2024. It sets the bar for what a genuinely [open-source model](https://geotoolbox.ai/glossary/open-source-ai) has to provide. The definition grants four freedoms, inherited from free software: the freedom to use the system for any purpose, to study how it works, to modify it, and to share it. To make those freedoms real, an open-source model has to release the "preferred form for making modifications." For an AI system, that means three things together: the model weights, the full code used to train and run it, and sufficiently detailed information about the training data. There is one nuance that trips people up. OSAID does not force a lab to publish the raw training dataset, which is often impossible to redistribute for copyright or privacy reasons. It requires detailed data information instead: where the data came from, how it was selected and filtered, and enough documentation that a skilled person could assemble a substantially equivalent dataset and retrain the model. Critics argue this is too soft and that true openness should require the data itself. That debate is still live. The practical test is reproducibility. If you have the weights, the training code, and the data information, an independent team can rebuild the model, inspect it for bias or backdoors, and verify what it claims to be. Very few models clear that bar. The clearest current example is OLMo from the Allen Institute for AI, now on its third generation: OLMo 3 ships its weights, full training code, checkpoints, and evaluation tooling under Apache 2.0, alongside the roughly 6-trillion-token Dolma 3 dataset it was pretrained on. It has company: EleutherAI's Pythia and LLM360's Amber and CrystalCoder are among the models the Open Source Initiative has pointed to as genuinely conforming. That is what open source actually looks like. ## The Openness Spectrum: Closed, Open-Weight, Open-Source Treating this as open versus closed hides the part that matters. Openness is a spectrum, and most of the interesting models sit in the middle.
![The AI openness spectrum from closed to open-weight to open-source, showing what you get at each tier: closed gives API access only, open-weight adds downloadable weights, open-source adds training code and data information.](/blog/open-weights-vs-open-source/openness-spectrum.png)
Openness is a spectrum. Most models people call "open" are open-weight, not open-source.
At one end are closed models like GPT-5.x, Claude, and Gemini. You reach them through an API. You never touch the weights, and you cannot run them yourself. In the middle are open-weight models: the weights are downloadable, but the recipe is not. At the far end are open-source models that meet OSAID, where the full stack is public.
TierWhat you getExample modelsTypical license
ClosedAPI access only; no weightsGPT-5.x, Claude, GeminiProprietary terms of service
Open-weightDownloadable weights; run and fine-tune; no training code or dataLlama, DeepSeek, Qwen, Gemma, gpt-oss, Kimi, InklingCustom or permissive, varies
Open-source (OSAID)Weights, training code, and data information; fully reproducibleOLMoOSI-approved (Apache 2.0)
There is a fourth label worth knowing: restricted weights. These are open-weight models released under terms that limit who can use them or how, often for regulated industries. The Open Source Initiative keeps a hard line here. In its view, any restriction on use or field of endeavor means a model is source-available or proprietary, not open source. Others prefer the spectrum framing, arguing that "usable but limited" deserves its own category rather than being lumped in with fully closed systems. ## Which AI Models Are Open Weight vs Open Source? (2026) Now the useful part: which real models are which. This is where the labels get sloppiest, because nearly every model below is called "open source" somewhere. Only one of them actually is. For ranked verdicts on which of these to use, see our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) list.
ModelOpen-weight or open-source?LicenseCommercial-use catch
Llama (Meta)Open-weight; frontier work has moved closedLlama Community License + Acceptable Use PolicyFree unless you exceed 700M monthly active users; use-case and regional bans apply
OLMo (Allen Institute for AI)Open-sourceApache 2.0None; full code and Dolma data released
DeepSeek (R1, V3-0324 and later)Open-weightMITNone; distillation and commercial use allowed
Qwen (Alibaba)Open-weight, up to a pointApache 2.0 (Qwen3 series, through Qwen3.6)None on the Apache 2.0 line (the Qwen3 series through Qwen3.6, resumed with Qwen3.8-27B in August 2026); the 3.7 generation is API-only. On August 12, 2026 Alibaba published Qwen3.8-Max weights (text-only variant) under a bespoke "Qwen3.8-Max" license, its first open-weight Max-class model
Kimi (Moonshot)Open-weightModified MIT (K2 family); custom Kimi K3 License (K3)Display or separate-agreement requirements above very large user, revenue, or MaaS thresholds
GLM-5.2 (Zhipu / z.ai)Open-weightMITNone
Inkling (Thinking Machines)Open-weightApache 2.0None; training data and code withheld
MistralOpen-weight (varies by model)Apache 2.0 or research/non-production license, by modelResearch-licensed models are not for commercial use
Gemma (Google)Open-weightApache 2.0 (Gemma 4)None on Gemma 4; earlier Gemma used custom Google terms
gpt-oss (OpenAI)Open-weightApache 2.0None
Read that top row carefully, because it is the one people get wrong most. Meta calls Llama open source in its own marketing, but [Wikipedia's summary](https://en.wikipedia.org/wiki/Llama_(language_model)) reflects the view of the definition-setting bodies: Llama's license discriminates against certain users and uses, which the Open Source Definition forbids, so it is not open source. The Free Software Foundation reached the same conclusion. Meta's [Community License](https://developer.meta.com/ai/llama4/license/) requires anyone above 700 million monthly active users to ask for permission, which Meta can refuse. The permissive open-weight releases are more generous. [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek) ships its recent models under the [MIT license](https://github.com/deepseek-ai/DeepSeek-R1), [Qwen](https://geotoolbox.ai/blog/what-is-qwen) ships most of its open line under [Apache 2.0](https://github.com/QwenLM/Qwen3) (through Qwen3.6, and again with Qwen3.8-27B in August 2026 after the 3.7 generation shipped no open weights), and [OpenAI's gpt-oss](https://github.com/openai/gpt-oss) models are Apache 2.0 as well. [Moonshot's Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai) K2-family models use a Modified MIT license that adds an attribution requirement at scale, though its newer K3 ships under a custom license with firmer conditions, and Zhipu's [GLM-5.2](https://geotoolbox.ai/blog/what-is-glm-5-2) is MIT too. Google flipped [Gemma 4 to Apache 2.0](https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/) in April 2026, dropping the custom terms its earlier models carried. In our experience helping brands track how AI systems cite them, the license is the detail teams skip and later regret, because the download page and the terms of use tell two different stories. Notice the pattern across the whole [field of Chinese open-weight models](https://geotoolbox.ai/blog/chinese-ai-models-compared) and beyond: permissive licenses are common, but open source, in the strict sense, is rare. ### "Announced Open-Weight" and "Downloadable" Are Not the Same Day There is a third gap underneath the two this article is about, and Moonshot's [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3) showed it plainly. K3 was announced on July 16, 2026, at 2.8T parameters with a million-token context window, and covered everywhere as an open-weight release. For its first eleven days it was not one: there was no K3 repository on Moonshot's Hugging Face org, and the model was reachable only through Moonshot's own apps and API. Moonshot published the weights on July 27, 2026, closing the gap, but the launch window is the point. "Open-weight" gets applied to a model at announcement, sometimes weeks before anyone outside the lab can download anything, and the press cycle that shapes people's mental model of a release happens entirely inside that window. The same discipline that applies to licenses applies here: check the model card, not the headline. K3 proved the point on the license too. Much of the coverage assumed it would inherit the Modified MIT terms of Moonshot's earlier models; instead it arrived under a new, custom Kimi K3 License. That license permits commercial use but is not open source: a model-as-a-service operator whose revenue passes $20 million over any 12 months must negotiate a separate agreement with Moonshot, and any product above 100 million monthly users or $20 million in monthly revenue must display "Kimi K3" in its interface. Expecting Modified MIT was reasonable, but it was not a fact, and the real terms only appeared on the model card. (For what those weights take to actually self-host, see [how to run Kimi K3 locally](https://geotoolbox.ai/blog/how-to-run-kimi-k3-locally).) ## Why the Difference Matters This is not a semantic argument. The gap between open weights and open source has three consequences that land on real projects. The first is legal. Downloading the weights does not give you the right to do anything you want with them. The license does. A custom open-weight license can cap you at a user threshold, ban entire fields of use, forbid using the model's outputs to train a competitor, or exclude whole regions. Meta's multimodal Llama models, for example, are not licensed to individuals or companies based in the European Union. The permissive open-source licenses, MIT and Apache 2.0, carry almost none of that; Thinking Machines' [Inkling](https://geotoolbox.ai/blog/inkling-ai), for instance, ships its weights under Apache 2.0. So "is it free to use commercially" has no general answer. It depends on which model and which release. The second is trust. With an open-weight model you inherit the creator's decisions about training data, filtering, and safety tuning, and you cannot fully audit any of them. That is a genuine liability in regulated sectors like finance or healthcare, where you may need to prove what a model was and was not trained on. Downloadable is also not the same as vetted. You cannot read a block of weights for a hidden backdoor the way you can read source code, and poisoned model files do surface on public hubs, so open weights still need scanning rather than blind trust. Only a reproducible, open-source model lets an outside team verify those claims in full. The third is control. Open weights let you fine-tune, which is cheap and practical. They do not let you retrain from scratch, which for a frontier model is wildly expensive and, without the data and code, impossible anyway. So your ability to truly change the model is bounded. Self-hosting is not the automatic saving people expect, either: renting the GPUs and keeping a large model running usually costs more than a metered API until you are serving very high volume. And an open-weight model still depends on its vendor. Terms can change, and a model can be [pulled or restricted](https://geotoolbox.ai/blog/fable-5-ban), as export controls have already shown, leaving teams scrambling for a replacement. The capability cost of choosing open is smaller than most people assume. [Epoch AI finds](https://epoch.ai/data-insights/open-weights-vs-closed-weights-models) that the best open-weight models trail the best closed models by only a few months, about three on average by its tracking. When the performance gap is that narrow, licensing and control become the real decision, not raw capability. ## Why Companies Release Weights but Call It "Open Source" If open weights are not open source, why do so many labs use the term anyway? Part of it is genuine, and part of it is marketing. The genuine part is that publishing full training data and code is hard and risky. The data often contains copyrighted or private material a company cannot legally redistribute. The training recipe is a trade secret worth hundreds of millions in compute and research. And a fully open pipeline makes it easier for someone to strip out the safety tuning. Releasing weights while holding back the rest is a real compromise between openness and those pressures. The marketing part is where it gets slippery. "Open source" carries goodwill that "open weights" does not, so labs reach for the stronger term. Critics call this openwashing: dressing up a weights-only release as something more transparent than it is. When Meta published its 2024 essay framing Llama as open source, the Open Source Initiative pushed back publicly, and the dispute became the clearest example of the gap between the label and the definition. There is even a linguistic argument that the word has drifted. In this view "open source" has fossilized into a loose synonym for "not locked behind an API," with the "source" part no longer meant literally. Whether you find that reasonable or misleading, it is why precise terms matter. ### Who Is Still Releasing Weights in 2026 The more interesting question this year is not what labs call their releases but which labs are still making them. Two of the archetypal open-weight vendors have quietly moved their best work behind an API. Meta launched [Muse Spark](https://about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/) in April 2026, its first proprietary closed-weight frontier model, available in private preview to select partners; Meta says it hopes to open-source future versions, without committing to a timeline. Alibaba did the same across two straight generations: Qwen3.7-Max in May and the Qwen3.7-Plus mid-tier in June were both API-only, and the downloadable Qwen line stopped at Qwen3.6 — a closed-flagship pattern that held right up until [Qwen3.8-Max](https://geotoolbox.ai/blog/qwen3-8-max) in August. OpenAI's gpt-oss weights have not been refreshed since their August 2025 release. For Meta and OpenAI's gpt-oss, the older weights stay up and nothing is formally discontinued, but the frontier moved closed. Alibaba is the one to watch, because it has now reversed its own pattern. The Qwen3.8-Max launch post committed to publishing the weights in the week of August 10, 2026, and on August 12 they landed in Alibaba's Hugging Face organization: a text-only variant under a bespoke "Qwen3.8-Max" license rather than Apache 2.0. A lab reopening its frontier, even partially and on custom terms, is the first real counterexample to the direction this section describes. Meanwhile the open-weight releases kept coming, and mostly from Chinese labs: [DeepSeek V4](https://geotoolbox.ai/blog/deepseek-v4) under MIT, GLM-5.2 under MIT, Tencent's Hunyuan Hy3 under Apache 2.0, and Kimi K3, whose weights went public on July 27 under its own custom license. The tempting conclusion is that open weights are becoming a Chinese specialty, and that overstates it. Google moved the other way, flipping Gemma 4 to Apache 2.0 in April 2026. Thinking Machines Lab released [Inkling](https://geotoolbox.ai/blog/inkling-ai) in July 2026 under Apache 2.0, a 975-billion-parameter mixture-of-experts model reported as the largest US open-weight release to date, though its training data and code are withheld, which makes it open-weight rather than open-source. The accurate version is narrower and more useful: the open-weight frontier is increasingly, but not exclusively, Chinese, and the labs with the most to lose commercially are the ones pulling back. Others are trying to fix the vocabulary rather than fight it. The [Open Weight Definition](https://www.forbes.com/sites/adrianbridgwater/2025/01/22/open-weight-definition-adds-balance-to-open-source-ai-integrity/), from the Open Source Alliance, gives weights-only releases their own honest standard instead of forcing them under the open-source banner. The practical takeaway for anyone choosing a model is simple: ignore the label on the announcement and read the license on the model card. That is the only thing that actually governs what you can do. ## What Open Weights vs Open Source Means for AI Visibility For most marketers this reads as an engineering debate, but it shapes where your brand does and does not get mentioned. The biggest answer engines, ChatGPT and Google's AI Overviews and Gemini and Copilot, still run on closed frontier models. Open weights widen the long tail underneath them. Anyone can run DeepSeek, Qwen, Kimi, Llama, or gpt-oss, so a growing set of assistants, search features, and [web-browsing agents](https://geotoolbox.ai/blog/how-does-ai-search-work) is built on downloadable models rather than a big lab's API. Perplexity's Sonar, for one, is fine-tuned from Meta's Llama. The lever that decides whether you get named is the same across all of them, open or closed. None of these systems let you touch the weights, but every one of them pulls in outside sources at answer time and picks which to cite. That evidence is the part you can influence, and open weights simply mean more systems are making the call. ## Frequently Asked Questions ### Is Llama open source or open weight? Open weight. Meta calls Llama open source, but its Community License restricts certain users and fields of use, which the Open Source Definition forbids, so the Open Source Initiative and the Free Software Foundation both say it is not open source. You can download and fine-tune the weights, but the training data and code are not released, and companies above 700 million monthly active users need a separate license from Meta. ### Is DeepSeek open source or open weight? Open weight. DeepSeek's R1 weights and its V3 releases from V3-0324 onward are published under a permissive MIT license, so commercial use and even distillation are allowed. (The original V3 base/chat weights predate that switch and carry a separate custom Model License.) But DeepSeek does not release its training data or full training code, so you cannot reproduce or fully audit the model. Very permissive, still not open source in the strict sense. ### Is ChatGPT open source? What about gpt-oss? ChatGPT and the GPT-5.x models behind it are closed. You only reach them through OpenAI's API or app. Separately, OpenAI released gpt-oss-120b and gpt-oss-20b as open-weight models under Apache 2.0, which you can download and run yourself. Those are open weight, not open source, because the training data and code stay private. ### Is Qwen open source? Partly. Qwen's open-weight models, the Qwen3 series through Qwen3.6, are released under Apache 2.0, so they are free for commercial use with no user cap. Older Qwen models used a custom community license with more conditions. But Alibaba's 3.7 generation stayed closed and API-only, from the Qwen3.7-Plus mid-tier to the Qwen3.7-Max flagship, so the open line and the frontier line had diverged. That changed on August 12, 2026, when Alibaba published the Qwen3.8-Max weights, its first open Max-class model, as a text-only variant under a bespoke "Qwen3.8-Max" license. That makes the flagship open-weight, not open source: no training code, no data disclosure, and a custom license rather than a permissive one. ### Can I use open-weight models commercially? Usually yes, but the license decides, not the download button. Permissive licenses like MIT and Apache 2.0 allow commercial use freely. Custom licenses may add user thresholds, attribution requirements, field-of-use bans, or regional exclusions. Always read the license on the model card before you build on it. ### What is the difference between open weights and open source in one sentence? Open weights give you the finished model to run and fine-tune, while open source gives you the model plus the training code and data information needed to rebuild and fully audit it. ## The Label Is Loose, the License Is Not The word "open" is doing a lot of unearned work in AI right now. What almost always sits behind it is open weights: downloadable and genuinely useful, but not reproducible and not free of strings. That is not a scandal, just a different thing, with real consequences for what you are allowed to do and what you can trust. So read the license on the model card, not the label on the announcement. If your buyers are getting answers from AI, being named in those answers is its own problem to solve, whichever models the engines happen to run on. geotoolbox's [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) shows which offsite sources the major AI engines cite when they answer questions in your market, and where your brand is missing from them, so you can work on the evidence these systems actually retrieve. ## Sources - Open Source Initiative - The Open Source AI Definition 1.0 (OSAID) - `opensource.org/ai/open-source-ai-definition` - Wikipedia - Llama (language model): open source status and license - `en.wikipedia.org/wiki/Llama_(language_model)` - Meta - Llama 4 Community License Agreement - `developer.meta.com/ai/llama4/license` - Google - Gemma 4 released under Apache 2.0 - `blog.google/innovation-and-ai/technology/developers-tools/gemma-4` - OpenAI - gpt-oss open-weight models, Apache 2.0 (GitHub) - `github.com/openai/gpt-oss` - DeepSeek-R1 - MIT License (GitHub) - `github.com/deepseek-ai/DeepSeek-R1` - Qwen3 - Apache 2.0 open-weight models (GitHub) - `github.com/QwenLM/Qwen3` - Qwen3.8-Max: A New Bar for Coding and Cowork (the Max-class open-weights commitment) - Qwen / Alibaba, August 3, 2026 - `qwen.ai/blog?id=qwen3.8` - Qwen models - Hugging Face (checked for the Qwen3.8-Max weights, August 7, 2026) - `huggingface.co/Qwen` - Epoch AI - Open-weight models lag the best closed models by about three months - `epoch.ai/data-insights/open-weights-vs-closed-weights-models` - Introducing Muse Spark (Meta Superintelligence Labs) - `about.fb.com/news/2026/04/introducing-muse-spark-meta-superintelligence-labs/` - Moonshot AI on Hugging Face (Kimi model cards and licenses) - `huggingface.co/moonshotai` - OLMo 3 (Allen Institute for AI) - `allenai.org/blog/olmo3` - Forbes - Open Weight Definition Adds Balance To Open Source AI Integrity - `forbes.com/sites/adrianbridgwater/2025/01/22/open-weight-definition-adds-balance-to-open-source-ai-integrity` --- ## What Is Claude AI? Anthropic's AI Assistant, Explained > What is Claude AI? Anthropic's assistant explained: the models (Haiku, Sonnet, Opus, Fable 5), pricing, limits, privacy, and what it means for your brand. - Canonical: https://geotoolbox.ai/blog/what-is-claude-ai - Published: 2026-07-03 · Updated: 2026-08-14 Claude AI is the chat assistant made by Anthropic, best known for careful, low-fluff answers and a reputation for handling long documents and code especially well. If you have heard the name next to ChatGPT and Gemini and want the plain version of what it is, who builds it, what it costs, and what it can actually do, this is it, current as of August 2026. We will also cover the parts most explainers skip: the usage limits people complain about, what happens to your data, and the new Fable 5 model that the US government pulled offline within days of launch, then let back a few weeks later. ## What Is Claude AI? **Claude is a family of [large language models](https://geotoolbox.ai/glossary/large-language-model) and an AI assistant built on top of them, made by Anthropic.** You talk to it the way you would text a sharp colleague: ask a question, paste a document, describe a task, and it answers in natural language. It is the same broad category of tool as ChatGPT and [Gemini](https://geotoolbox.ai/blog/what-is-gemini), not a different species. You can use Claude free at [claude.ai](https://claude.ai), in the iOS, Android, and desktop apps, or through Anthropic's API, which is how other products plug the same models in behind the scenes ([Microsoft's Copilot](https://geotoolbox.ai/blog/what-is-copilot), for instance, lets you choose Claude as its model). The models are closed and proprietary, so Claude only runs through Anthropic or partners who license it. Claude works mainly through text. It reads documents and images, writes and analyzes code, and on the mobile app can hold a spoken conversation, but it does not generate images or video the way ChatGPT and Gemini do. Calling it a chatbot undersells it, though. It can work through a long report, keep a multi-step task on track, build small interactive tools in a panel beside the chat, and search the live web when a question needs current facts. The rest of this guide walks through each of those pieces. ## Who Makes Claude? Meet Anthropic Claude comes from [Anthropic](https://www.ibm.com/think/topics/claude-ai), an AI lab founded in 2021 by a group of former OpenAI staff, including the siblings Dario Amodei, its CEO, and Daniela Amodei, its president. They left before OpenAI released ChatGPT, in part to put AI safety and interpretability (the study of what is actually happening inside these models) at the center of the work. Anthropic is registered as a public benefit corporation, which means its charter formally weighs public interest alongside profit. This answers one of the most common questions about Claude: who owns it? **Anthropic owns and runs Claude.** Amazon is its largest outside investor, with commitments reported in the tens of billions, and Google holds a minority stake, but neither company owns or controls Anthropic, and neither one decides how Claude behaves. The confusion is understandable, because both Amazon and Google also sell access to Claude through their own clouds, but the model is Anthropic's. That safety-first origin is not just branding. It shapes how Claude is trained, why it sometimes refuses or hedges where other assistants charge ahead, and the unusual situation it found itself in this June, when its most capable model was switched off by a government order. More on both shortly. ## How Claude Works, and What Makes It Different Under the hood, Claude is a transformer that predicts the next chunk of text one piece at a time, trained on a huge amount of writing and code until patterns of grammar, fact, and reasoning settle into its billions of internal numbers. That part is true of ChatGPT and Gemini too, so it is not where Claude stands apart. The difference is how Anthropic shapes its behavior. Most assistants are tuned mainly with human feedback, where people rate answers and the model learns to prefer the better-rated ones. Anthropic adds a method it calls **Constitutional AI**: Claude is trained against a written set of principles, [Claude's Constitution](https://www.anthropic.com/constitution), and learns to critique and revise its own answers toward being helpful, honest, and harmless. The practical result is the trait people notice first, a model that tends to be careful, explains its reasoning, and pushes back instead of confidently inventing an answer. Newer Claude models are also hybrid reasoning models. By default they reply right away, but you can switch on extended thinking for hard problems, where the model works through the steps before answering. That trades speed for accuracy on math, analysis, and complex coding. This is the short version. If you want the honest, deeper walkthrough, including what researchers found by looking inside the model and why even Claude's own explanation of its reasoning can be a guess, read our companion piece on [how Claude works](https://geotoolbox.ai/blog/how-does-claude-work). ## The Claude Model Lineup: Haiku, Sonnet, Opus, and Fable
![Four cards showing Claude's model lineup from Haiku 4.5 up to Fable 5.](/blog/what-is-claude-ai/claude-model-lineup.png)
The Claude lineup in August 2026, from the fast Haiku tier up to the new Fable frontier tier.
Claude is not one model but a lineup, named after forms of writing in rough order of size and power. Haiku is the short, fast one. Sonnet is the medium, balanced one. Opus is the large, most capable one. In June 2026, Anthropic added a new top tier above Opus, beginning with a model called Fable 5. Here is where things stand.
Model (August 2026)Speed and costBest forStatus
Claude Haiku 4.5Fastest, cheapestHigh-volume, simple tasks and quick answersAvailable
Claude Sonnet 5BalancedThe default for most work: writing, analysis, everyday codingAvailable (new default, launched June 30, 2026)
Claude Opus 5Most capable Opus tier, slower, priciestHard reasoning, deep research, complex codingAvailable (launched July 24, 2026, replacing Opus 4.8 at the same price)
Claude Opus 4.8Same price as Opus 5Hard reasoning, deep research, complex codingAvailable as a legacy model
Claude Fable 5New tier above OpusFrontier performance across benchmarksAvailable again (restored globally July 1, 2026)
Claude Mythos 5Same model as Fable, some safeguards liftedApproved customers only (Project Glasswing)Restored for a set of US organizations (June 26, 2026 approval)
[Claude Opus 4.8](https://www.anthropic.com/news/claude-opus-4-8) shipped on May 28, 2026, and was the most advanced model in the classic Opus line until [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5) succeeded it on July 24, 2026 at the same price. Fable 5 launched two weeks later as the first public model in a new, more powerful "Mythos-class" tier, went dark within days under a US government order, and was restored on July 1, 2026 (more on that next). **Which one should you pick?** If you are unsure, start with Sonnet, the everyday workhorse. Reach for Haiku when you want speed and low cost on routine tasks, and switch to Opus for genuinely hard problems. On the paid plans you can change models per conversation, so the choice is never permanent. Because version numbers move every few weeks, treat this table as an August 2026 snapshot. ## Wait, Didn't Claude Just Get Banned? Not Claude as a whole, but its two most powerful models. This is the news driving a lot of the "is Claude banned" searches, so here is the short, dated version. Fable 5 and its less-restricted sibling Mythos 5 went live on June 9, 2026. Three days later, on June 12, the US government issued a national-security [export-control directive](https://www.anthropic.com/news/fable-mythos-access) ordering Anthropic to suspend access to both models for any foreign national, anywhere, including Anthropic's own foreign-national employees. Complying with a rule that broad was all-or-nothing, so Anthropic switched Fable 5 and Mythos 5 off for everyone. Every other Claude model, including Opus 4.8, kept working normally. The trigger, as reported, was a claimed method of bypassing Fable 5's safety guardrails. Anthropic has said it believed the issue was narrow rather than a universal jailbreak, and outside coverage has questioned the order's footing, with [TechCrunch reporting](https://techcrunch.com/2026/06/15/the-us-governments-anthropic-models-ban-was-never-about-an-ai-jailbreak/) that the ban "was never about an AI jailbreak." The story has since resolved. On June 30, 2026, the Commerce Department lifted the export controls, and Anthropic [redeployed Fable 5](https://www.anthropic.com/news/redeploying-fable-5) globally on July 1, 2026, including in Europe, the API, and Claude Code. Mythos 5 returned for a set of US organizations following a June 26 government approval. It moved fast, so treat this as an August 2026 snapshot and check Anthropic's announcements for the current state. For the full breakdown of what Fable 5 and Mythos 5 are and why they were pulled and restored, see our [Fable 5 and Mythos 5 explainer](https://geotoolbox.ai/blog/fable-5-ban). ## Is Claude AI Free? Plans and Pricing So how much does Claude AI cost? Less than you might expect to start, because there is a real free tier. You can sign up at claude.ai with an email and phone number and use Claude on the web and in the apps. The catch is usage: the free plan has the lowest caps, which reset on a rolling basis roughly every five hours.
PlanPrice (August 2026)What you get
Free$0Claude on web and apps, web search, the lowest usage limits
Pro$20/monthAround five times the free usage, access to Opus, Projects, Claude Code
Max 5x$100/monthFive times Pro usage, priority access, higher Claude Code limits
Max 20x$200/monthTwenty times Pro usage, the highest individual limits
Team and EnterprisePer seat or customShared workspaces, admin controls, single sign-on, larger context
Prices and limits here are an August 2026 snapshot, so check claude.ai for the current numbers. Paid plans give you more usage and fuller access to the most capable Opus model, while the free tier defaults to a fast Sonnet-class model and has the tightest limits. The current Opus and Sonnet models can work with up to a 1-million-token [context window](https://geotoolbox.ai/glossary/context-window), roughly a long book and then some in a single conversation, though how much of that the chat app exposes depends on your plan (Haiku tops out lower, at 200,000 tokens). Developers pay separately through the API, priced per million tokens of input and output (Sonnet 5 runs $2 in and $10 out, for example; Anthropic announced that rate as introductory but made it permanent in August 2026, cancelling the planned rise to $3/$15). **Is Pro worth it?** If you hit the free limits often enough to be interrupted, or you want priority access to Opus, $20 is the usual answer. If you only dip in occasionally, the free tier is genuinely usable. ## The Usage Limits Everyone Complains About If you read about Claude before trying it, the loudest complaint you will find is about running out of usage, sometimes even on a paid plan. This is the part most explainers gloss over, so here is how it actually works. Claude limits you on a **rolling window**, commonly described as about five hours, and paid tiers add weekly caps on top. The key thing to understand is that not all messages cost the same. A short question to Haiku barely moves the needle; a long prompt to Opus, especially one carrying a big document, eats far more of your allowance. So two people on the same plan can hit the wall at very different points depending on which model they use and how much they paste in. Limits also tighten during peak hours, when overall demand on Anthropic's systems is highest. There is a second, quieter limit: the context window. When a single conversation grows past the window your plan allows, the app may start summarizing or dropping the earliest parts to keep going, which is why a very long thread can begin to feel like it is forgetting details from the top. The practical fix is to match the model to the job. Route routine drafting, summarizing, and quick lookups to Sonnet or Haiku, and save Opus for the genuinely hard problems. Start a fresh chat for a new topic instead of letting one thread run for hours, and you will stretch your allowance a lot further than fighting a single 50-message conversation will. ## Does Claude Train on Your Conversations? The honest answer is "only if you let it," and the default changed in 2025, so it is worth getting right rather than trusting a flat yes or no. For free, Pro, and Max accounts, Anthropic asks whether you want to share your chats to [help improve Claude](https://www.anthropic.com/news/updates-to-our-consumer-terms). If you leave that setting on, your conversations can be used to train future models and are kept for up to five years. If you turn it off, your chats are not used for training and fall under a much shorter retention window of around 30 days. You can change the setting at any time in your privacy controls, and deleting a conversation or your account removes that data from future training. Business usage is treated differently. Traffic through Claude for Work, Enterprise, Education, and the API is not used to train models by default, which is part of why companies are comfortable putting Claude into their workflows. So the myth cuts both ways. Claude does not silently hoover up everything you type, and it is also not a guaranteed privacy vault. The control is a toggle, and the sensible move is to open your settings and decide for yourself rather than assume either extreme. ## What Can Claude Do, and Do You Need to Code? Claude's strengths cluster around a few things: writing and editing that reads naturally, working through long documents without losing the thread, structured reasoning, and coding, where it is widely rated near the top. It can also run multi-step tasks fairly autonomously, the "agentic" behavior every AI company is chasing. **Can Claude access the internet?** Yes. This is one of the most stubborn outdated beliefs about it. Claude has built-in web search and will pull live pages when a question needs current information, then cite the sources it used. Anthropic also runs a web crawler, [ClaudeBot](https://geotoolbox.ai/glossary/claudebot), though that one mainly gathers training data; the live search built into your chat fetches pages in the moment. There is still a knowledge cutoff baked into the model from training, but web search is how Claude reaches past it for anything recent. A point of confusion worth clearing up: **the Claude chat app and Claude Code are two different things.** The app at claude.ai is general-purpose and built for everyone, writers, marketers, students, analysts, no coding required. [Claude Code](https://geotoolbox.ai/blog/what-is-claude-code) is a separate tool for software engineers that runs in a terminal and edits real codebases. If you have read intimidating articles about Claude writing software and wondered whether Claude is "for developers," the answer is that the chatbot most people mean is for everyone; Claude Code is the specialist add-on. One real limit to keep in mind: Claude reads images and documents, but it does not generate photographic images or video the way ChatGPT and Gemini do. (Anthropic's separate Claude Design preview can build editable layouts like slides and prototypes, but not pixel images.) For a finished illustration you still reach for a dedicated image tool. Claude's lane is text, reasoning, and code, and it is deliberately deep rather than wide. ## Claude vs ChatGPT vs Gemini, the Short Version People rarely evaluate Claude in a vacuum; the real question is usually how it stacks up against ChatGPT and Gemini. The honest answer is that they are close enough that the "best" one depends on the task, not a scoreboard.
AssistantMakerLeans best at
ClaudeAnthropicLong-document work, natural writing, careful reasoning, coding
ChatGPTOpenAIImage and video generation, voice, advanced reasoning, custom GPTs, the widest ecosystem
GeminiGoogleNative multimodality, long-context retrieval, deep Google Search and Workspace integration
In plain terms: Claude is known for natural-sounding writing and a careful, low-fluff tone; ChatGPT covers the most ground and is the one to reach for if you also want image generation, voice, and the largest plugin ecosystem; and Gemini's edge is native multimodality and living inside Google's products and index. On most individual tasks the latest models are close, which is why plenty of people keep two open and switch by job. If you are weighing the open-weight challengers too, [Kimi K2](https://geotoolbox.ai/blog/what-is-kimi-ai), [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek), and [Zhipu's GLM 5.2](https://geotoolbox.ai/blog/what-is-glm-5-2) now match the big closed models on some tasks at a fraction of the cost. We go deep on the head-to-heads elsewhere, including where each one genuinely wins and where the marketing oversells it, in [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt), [Claude vs Gemini](https://geotoolbox.ai/blog/claude-vs-gemini), and [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude). If you are choosing one assistant to live in, those comparisons are the place to settle it. For the wider field beyond Claude, our roundup of the [best ChatGPT alternatives](https://geotoolbox.ai/blog/chatgpt-alternatives) sorts every major assistant by the job you need it for. ## What Claude Means for Your Brand Here is the part most "what is Claude AI" explainers never reach, and the reason it matters if you run a website or a brand. When Claude searches the web, it does more than answer; like ChatGPT, Gemini, and Google's AI Overviews, it names the sources it leaned on. The result at the top of a search like that is now often an AI summary, not ten blue links. So a new question sits underneath the old SEO one. When a customer asks Claude "what's the best tool for X" or "is [your company] any good," does your brand come up, and is what Claude says about it correct? Claude draws on two places: what it absorbed during training, and what its [web search and crawler](https://geotoolbox.ai/blog/how-does-ai-search-work) pull live. Both are influenced by how clearly and consistently your business is represented across the web. In our experience at geotoolbox, the brands that show up well in AI answers are rarely the ones with the slickest homepage; they are the ones whose facts are consistent, well-structured, and easy for a model to retrieve and trust. That is a fixable, measurable problem, which is the whole reason we build tooling around it: see what AI assistants actually say about you and where you are missing. Our guides on [getting cited in Claude](https://geotoolbox.ai/blog/claude-seo) and [tracking your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) go deeper. Understanding what Claude is matters most when you stop asking "what can it do for me" and start asking "what does it already say about me." You can answer that one today. Run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether Claude and the other AI engines can find and read your site, and where the gaps are. ## Frequently Asked Questions ### Is Claude AI free? Yes. There is a genuine free tier at claude.ai that needs only an email and phone number, no credit card. It has the lowest usage limits, which reset on a rolling basis every few hours. If you hit those limits regularly, [Claude Pro](https://geotoolbox.ai/blog/claude-pricing) is $20 a month for roughly five times the usage and priority access to the top Opus model. ### Who owns Claude AI? Anthropic owns and operates Claude. Amazon is its largest outside investor and Google holds a minority stake, but neither company owns or controls Anthropic, and neither decides how Claude behaves. Anthropic was founded in 2021 by former OpenAI staff, including siblings Dario and Daniela Amodei. ### Is Claude better than ChatGPT? It depends on the task rather than a single winner. Claude is often praised for natural writing, long-document work, and a careful tone; ChatGPT is strong on reasoning, images, voice, and its wider ecosystem; and on most tasks the latest models are close. Many people use both and switch depending on the job. ### Can Claude access the internet? Yes. Claude has built-in web search and will pull live pages when a question needs current information, then cite the sources it used. It still has a training knowledge cutoff, but web search is how it reaches past that for recent facts. The belief that Claude cannot browse the web is out of date. ### Does Claude train on my conversations? Only if you allow it. Free, Pro, and Max accounts have a "help improve Claude" setting; leave it on and your chats can be used for training and kept up to five years, turn it off and they are not used for training and kept around 30 days. Business plans like Enterprise and the API are not used for training by default. ### Is Claude AI safe? Safety is Anthropic's founding pitch. Claude is trained with Constitutional AI, a written set of principles meant to keep it helpful, honest, and harmless, which is why it tends to refuse clearly harmful requests and hedge when it is unsure. No AI is flawless and Claude can still be wrong, but for everyday use it is as safe as the other major assistants. On data specifically, see the privacy note above and check your settings. ### Which Claude model should I use? Start with Sonnet, the balanced default that handles most work well. Use Haiku when you want speed and low cost on simple, high-volume tasks, and switch to Opus for genuinely hard reasoning, research, or coding. On paid plans you can change the model per conversation. ### What is the difference between Claude and Claude Code? The Claude app at claude.ai is the general-purpose assistant most people mean; it needs no coding skills and suits writing, research, and analysis. Claude Code is a separate tool for software engineers that runs in a terminal and edits real codebases. You do not need Claude Code to use Claude. ### Is Claude AI banned? No, everyday Claude works normally. In June 2026 the US government ordered Anthropic to suspend its two most powerful new models, Fable 5 and Mythos 5, for foreign nationals under export-control rules, so Anthropic disabled them for everyone to comply. All other models, including Opus 4.8, stayed available. The export controls were lifted on June 30, 2026, and Anthropic restored Fable 5 globally on July 1, 2026 (with Mythos 5 restored for certain US organizations). ## Sources - What Is Claude AI? - IBM Think, updated 2026 - `ibm.com/think/topics/claude-ai` - Claude's Constitution - Anthropic - `anthropic.com/constitution` - Introducing Claude Opus 4.8 - Anthropic, May 2026 - `anthropic.com/news/claude-opus-4-8` - Claude Fable 5 and Claude Mythos 5 - Anthropic, June 2026 - `anthropic.com/news/claude-fable-5-mythos-5` - Statement on the US government directive to suspend access to Fable 5 and Mythos 5 - Anthropic, June 2026 - `anthropic.com/news/fable-mythos-access` - Redeploying Fable 5 - Anthropic, July 2026 - `anthropic.com/news/redeploying-fable-5` - Introducing Claude Sonnet 5 - Anthropic, June 2026 - `anthropic.com/news/claude-sonnet-5` - The US government's Anthropic models ban was never about an AI jailbreak - TechCrunch, June 2026 - `techcrunch.com/2026/06/15/the-us-governments-anthropic-models-ban-was-never-about-an-ai-jailbreak` - Updates to our Consumer Terms and Privacy Policy - Anthropic, 2025 - `anthropic.com/news/updates-to-our-consumer-terms` - What is the Pro plan? - Anthropic Support - `support.claude.com/en/articles/8325606-what-is-the-pro-plan` --- ## ChatGPT Pricing August 2026: Plans, API Cost, and Is It Worth It? > Every ChatGPT price for 2026: Free, Go, Plus, the two Pro tiers ($100 and $200), Business, Enterprise, and API token costs, plus which plan is worth it. - Canonical: https://geotoolbox.ai/blog/chatgpt-pricing - Published: 2026-07-02 · Updated: 2026-08-18 ChatGPT costs anywhere from nothing to $200 a month, and that is before you touch the API. The consumer plans run Free at $0, Go at $8, Plus at $20, and two separate Pro tiers at $100 and $200. Teams pay per seat: Business is $20 to $25 per user, and Enterprise is a custom quote. Developers pay the OpenAI API by the token, on a completely separate bill. ChatGPT pricing moved more than once in 2026. Go went global in January, a second Pro tier appeared in April, OpenAI cut Business pricing that same month, the GPT-5.6 models (Sol, Terra, Luna) previewed in June and became generally available on July 9, OpenAI cut GPT-5.6 API prices on July 30, and on August 6 it made GPT-5.6 Luna the default model on Free and Go with uncapped text chats rolling out from mid-August. Most pricing guides you will find still quote a tier list that has changed. Below is every current price, reconciled and dated August 2026, plus the question the numbers exist to answer: which plan, if any, you should actually pay for. All prices here are US list prices, and they run lower in many markets (Go launched in India first for exactly that reason). A ChatGPT subscription and the OpenAI API are also separate products with separate billing, so paying for one does not give you the other. ## How Much Does ChatGPT Cost? Every Plan at a Glance Here is the whole consumer and team lineup in one place, at [US prices from OpenAI](https://openai.com/chatgpt/pricing/) as of July 2026. Subscription prices were untouched by both the July 30 API cut and the August 6 free-tier change.
PlanPrice (US)What it isBest for
Free$0/moGPT-5.6 Luna by default, uncapped text chats rolling out from mid-August, ads in some marketsCasual, occasional use
Go$8/moSame default model as Free, higher file and image limits, still ad-supported, no advanced toolsLight daily use on a budget
Plus$20/moFull tool suite, Deep Research, ad-freeMost working professionals
Pro (mid)$100/mo5x Plus limits, heavier Codex and research usePower users who hit Plus caps
Pro$200/mo~20x Plus limits, largest context OpenAI offers an individual, full SoraHeavy researchers and builders
Business$20/user (annual) or $25/user (monthly)Team controls, security, data excluded from trainingTeams of 2 to ~149
EnterpriseCustomData residency, SLA, dedicated capacityLarge or regulated organizations
APIPay per tokenProgrammatic access to the models, no chat appDevelopers and products
There are now two tiers that both say "Pro" in your account, at $100 and $200, and they are not the same plan. And the API in that last row is a separate product, priced per million tokens rather than per month. ![ChatGPT US pricing as of July 2026, cheapest to priciest: Free $0, Go $8, Plus $20, Pro $100, Pro $200, plus Business at $20 to $25 per user with a 2-seat minimum, Enterprise custom pricing, and a pay-per-token API.](/blog/chatgpt-pricing/chatgpt-plans-july-2026.png) ## ChatGPT Free: What $0 Gets You (and the New Ads) Free is a real product, not a locked demo, and it got substantially better on August 6, 2026. The default model on Free is now **GPT-5.6 Luna**, the fast variant of the current generation, replacing GPT-5.5 Instant. Starting the week of August 10, 2026, OpenAI is also rolling out uncapped text conversations for Free and Go and a **Think** button that lets Luna spend longer on a harder question. The rollout is staged, so if you still see a cap your account has not been switched over yet. You still get image generation, voice, and file uploads. Read "uncapped" narrowly, though, because it is easy to oversell. It covers text chats only, subject to what OpenAI calls abuse guardrails; files, images, voice, and image generation keep their own separate limits, and those still bite. What Free still does not get at any volume is **GPT-5.6 Sol**, the flagship, which starts at Plus. Terra, the balanced variant, does reach Free and Go, but only inside Codex, never in the main chat window. Almost no guide makes either distinction. The full Deep Research tool is not part of Free, though the tier does include five lightweight research runs a month. The change worth knowing about in 2026: since February 9, the Free tier shows [ads at the bottom of responses](https://geotoolbox.ai/blog/chatgpt-ads). The US pilot has since expanded to Canada, the UK, Australia, New Zealand, Japan, South Korea, and, as of an August 11, 2026 update, Mexico and Brazil. OpenAI emailed Free and Go users in the EEA and Switzerland on August 15, 2026 confirming ads are coming there too, "later this month," with personalization off by default (ads there initially target on conversation topic, location, device, time of day, and language rather than a user profile) until users opt in. Ads run on Free and Go whether you are logged in or signed out, while every paid tier from Plus up stays ad-free. But you do not have to pay to escape them: Free and Go users can switch to an ad-free mode that trades ads for fewer messages a day and no access to tools like image generation, and a Temporary Chat shows none. OpenAI labels the ads, limits personalized targeting to logged-in adults, and says they do not influence answers, but they are there. Free conversations can also be used to train OpenAI's models unless you turn that off in your data controls. For occasional use, drafting an email, asking a quick question, testing what the model can do, Free is genuinely enough, and the August change makes it more so. It also moves the wall: you no longer run out of conversation first, you run out of features. No Sol, no full Deep Research, and separate caps on file uploads and image generation. If you are hitting those and also the smaller [context window](https://geotoolbox.ai/glossary/context-window) on longer documents, that is the signal you have outgrown Free. ## ChatGPT Go ($8/Month): The Budget Tier, and When It's a Trap Go sits between Free and Plus at $8 a month. It raises your upload and image limits above Free and extends memory, but it deliberately leaves out the tools that make ChatGPT a work tool: no Deep Research, no Sora, no Agent Mode. The August 2026 change narrowed the gap further and is worth pricing in: Free and Go now run the same default model, GPT-5.6 Luna, and both are getting uncapped text chats and the Think button as the staged rollout reaches accounts. What $8 buys you is headroom on files, images and memory, not on conversation. OpenAI launched it in India in August 2025, then rolled it out to more than 170 countries by January 2026, where it quickly became the company's fastest-growing plan. The $8 price explains the growth better than the feature set does. The catch is the ads, and where you live decides whether it applies. When they arrived in February 2026, Go was included alongside Free, so in the markets where the pilot has landed you are paying $8 a month and still seeing ads, while giving up every advanced feature. The same ad-free mode Free has is available, but it trades the ads for fewer daily messages and no image generation. In the EU, ads have not launched yet, though OpenAI confirmed on August 15, 2026 that they are coming to the EEA and Switzerland later that month; until they do, that argument does not hold and Go has to be judged purely on what it leaves out. The step up to Plus is only $12 more, and it buys the full tool suite with no ads at all. If you are choosing between Go and Plus, the honest math is that Go rarely wins. Either you need so little that Free covers it, or you are doing real work, in which case the extra $12 for Plus changes what the tool can do, not just how much of it you get. Go makes the most sense in markets where $8 versus $20 is a meaningful gap, which is exactly where OpenAI has pushed it hardest. ## ChatGPT Plus ($20/Month): The Default for Most People Plus is the plan most people should start with, and it has held at $20 a month since February 2023, three and a half years, while the feature set kept growing. That alone makes it one of the steadier deals in AI subscriptions. What you get: **GPT-5.6 Sol** as your default model in a standard chat, with a selectable effort level, plus Deep Research for multi-step reports, limited Sora video, Codex for coding, Agent Mode for multi-step tasks, Projects and custom GPTs, and no ads. Two limits on that model line are worth being exact about. Terra never appears in a standard ChatGPT conversation on any plan, including this one; you reach it in ChatGPT Work, in Codex, or through the API. Luna does appear in a standard chat, but as the default on Free and Go, not as something a Plus subscriber selects. And the fourth model, **Sol Pro**, is not a Plus feature: it is reserved for the Pro, Business and Enterprise plans, which also add the Extra High effort level. The limit Plus users bump into first is Deep Research: on the last published breakdown, 10 full runs a month plus 15 lightweight ones, after which it switches to the lighter version. For most jobs that is plenty. For anyone running several deep reports a day, it runs out fast, and that is the specific pressure that pushes people toward a Pro tier. A note on per-window message caps, which the August change has partly overtaken. OpenAI used not to publish them at all, so for a long time any exact figure was guesswork. Then its help centre stated that Plus and Go users could send up to 160 messages with GPT-5.5 Instant every three hours, after which chats fell back to GPT-5.5 Instant mini, with 10 Thinking messages every five hours on Go. Those numbers describe the GPT-5.5 Instant era. From the week of August 10, 2026 OpenAI began removing the text-chat rate limit on Free and Go, so as that staged rollout reaches your account the 160-message figure no longer binds Go at all. OpenAI has not restated a per-window number for Plus, and reasoning limits have never been published, so treat any exact Plus figure you are quoted as an estimate. Is $20 worth it? The useful test is not a general yes or no, it is your own week. Use Plus normally for seven days and count how often you actually hit a wall, on message limits, on Deep Research, on context length. If you rarely hit one, Plus is right and you do not need to spend more. If you hit walls daily, that is real evidence for a Pro tier rather than a hunch. At the same $20, Plus competes directly with Claude Pro and Gemini's mid plan, so if you are weighing them, the [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) breakdown covers where each one pulls ahead. ## The Two ChatGPT Pro Tiers: $100 vs $200 (Yes, Both Exist) This is where most pricing guides go wrong, because they were written before April 2026. ChatGPT now has two tiers that both display "Pro" in your account: one at $100 a month and one at $200. They are separate plans, not the same plan billed two ways. OpenAI [added the $100 tier](https://www.macrumors.com/2026/04/09/openai-pro-subscription-tiers/) on April 9, 2026, to fill the awkward gap between $20 Plus and $200 Pro. Its lever is volume: roughly five times the Plus limits, a higher Deep Research allotment that OpenAI does not quantify publicly, and heavier Codex access, aimed at developers who kept hitting Plus caps but could not justify $200. OpenAI's own help centre says both Pro tiers "include the same core capabilities" and that the difference is the usage allowance, so treat claims that the $100 tier is missing specific tools with care; the company does not publish such a list. It also lines up against Anthropic's $100 Claude Max tier, which is almost certainly why it exists. The $200 tier is the ceiling for individual plans. It pushes limits to roughly 20 times Plus with Deep Research around 250 runs a month, adds full Sora video generation, and opens up the largest context window ChatGPT offers an individual: 128K tokens on the Instant models and 400K on the reasoning models, per OpenAI's own plan comparison. You will see "one million tokens" quoted widely for this tier. OpenAI does not publish that figure on any plan page, and it does not match the published per-plan numbers. The API is a separate product with its own, larger window ([1,050,000 tokens on GPT-5.6](https://developers.openai.com/api/docs/models/gpt-5.6-sol)), and no consumer plan exposes it. What actually separates the $200 tier from the $100 one is usage volume, which is exactly how OpenAI describes it. If you routinely feed the model large document sets or run long agentic sessions, that is what you are paying for.
FeaturePlus ($20)Pro ($100)Pro ($200)
Message limitsStandard~5x Plus~20x Plus
Deep Research / month10 full + 15 lightweightNot published~250
Codex accessStandardExpandedMaximum
Context window54K Instant / 256K reasoningSame as Plus128K Instant / 400K reasoning
Sora videoLimitedLimitedFull
Before you subscribe to "Pro," check the exact plan label and price in your billing settings. Because two tiers share the name, it is easy to click into the $200 plan when you meant the $100 one. The two do very different things to your bill. ## ChatGPT Business and Enterprise: Team Pricing Once you are buying for a group, the plans change shape. ChatGPT Business, which OpenAI renamed from "Team," costs $20 per user per month billed annually or $25 billed monthly after a price cut in April 2026, and it requires a minimum of two seats. That two-seat floor matters: there is no solo Business plan, so the practical entry point is about $40 a month. What you get for it is the team layer that individual plans lack, SOC 2 Type II compliance, SAML single sign-on, an admin console, dozens of connectors to tools like Slack and Google Drive, and, most importantly for client work, conversations excluded from model training by default. Enterprise is the custom-quote tier, generally aimed at organizations of around 150 seats and up. The per-seat price is negotiated and can run higher than Business depending on the contract, but you are paying for things Business does not include: data residency across multiple regions, a real SLA, dedicated capacity, audit logs, and custom deployment. For a regulated company, the data residency alone is often the deciding line.
BusinessEnterprise
Price$20/user (annual), $25 (monthly)Custom quote
Minimum seats2~150
Data excluded from trainingYesYes
Data residencyNoYes (multi-region)
SLA and dedicated supportStandardYes
The decision between them is less about price than about compliance. If your team is under 150 people and does not need data residency or a signed SLA, Business does the job. If regulators, not budgets, are driving the conversation, that is Enterprise territory. For a side-by-side with the other big per-seat team assistant, our [Microsoft Copilot pricing](https://geotoolbox.ai/blog/copilot-pricing) guide covers how the seat math compares. One plan this lineup leaves out on purpose is ChatGPT Edu, OpenAI's tier for universities and schools, which is arranged institutionally through the school rather than bought per seat. ## ChatGPT API Pricing: Pay per Token (Including GPT-5.6) The API is a different product with a different meter. There is no monthly subscription and no message cap. You pay per token, split into input (what you send) and output (what the model returns), and priced per million tokens. A [ChatGPT subscription](https://geotoolbox.ai/blog/how-does-chatgpt-work) does not include API access, and API credit does not give you the chat app. Here are the [current published rates](https://developers.openai.com/api/docs/pricing) for the main GPT-5 models, per million tokens.
ModelInput (per 1M)Output (per 1M)Best for
GPT-5.6 Sol (flagship)$5.00$30.00Hardest coding, agents, research, security
GPT-5.6 Terra (balanced)$2.00$12.00Everyday work at GPT-5.5-class quality
GPT-5.6 Luna (fast)$0.20$1.20High-volume, latency-sensitive tasks
GPT-5.4 Nano$0.20$1.25High-volume, simple tasks
GPT-5.4 Mini$0.75$4.50The common default
GPT-5.4 (prior gen)$2.50$15.00Prior-generation balanced
GPT-5.5 (prior flagship)$5.00$30.00Prior-generation reasoning
GPT-5.4 Pro / GPT-5.5 Pro$30.00$180.00Hardest problems (prior gen)
Two discounts cut those numbers hard. Cached input, where you reuse the same context across calls, drops as much as 90% off the input rate. The Batch API takes 50% off both input and output if you can accept a job running asynchronously with up to a 24-hour turnaround. Neither works for a real-time chat feature, but both matter a lot for bulk processing. Google's API applies the same 90% cached and 50% batch discounts, which our [Gemini API pricing](https://geotoolbox.ai/blog/gemini-api-pricing) guide works through in detail. On the newest models: OpenAI [previewed](https://openai.com/index/previewing-gpt-5-6-sol/) the **GPT-5.6** family, code-named Sol, Terra, and Luna, in late June 2026 and made it [generally available on July 9](https://geotoolbox.ai/blog/gpt-5-6) across ChatGPT, Codex, and the API. It is now the current flagship line, superseding GPT-5.5. Prices held from preview into general availability, then moved. On July 30, 2026, OpenAI cut GPT-5.6 Luna by 80% and Terra by 20%: Luna went from $1/$6 to $0.20/$1.20 per million tokens, and Terra from $2.50/$15 to $2.00/$12. Sol stayed at $5/$30. Subscription prices and quotas were untouched, but Terra and Luna now consume less of your allowance inside Codex. The cut reset the ladder rather than preserving it: Terra now undercuts the prior-generation GPT-5.4 it was originally priced to match, and Luna's $1.20 output rate is below GPT-5.4 Nano's $1.25. There is also a long-context tier most guides omit, priced higher: Sol at $10/$45, Terra at $4/$18, Luna at $0.40/$1.80. If your workload is API-based, GPT-5.6 is now what most builders should default to; GPT-5.5 and the GPT-5.4 line remain available as prior-generation options. ## Subscription vs API: Which Is Actually Cheaper? This is the question the pricing pages never answer, and the honest reply is that they meter completely different things. A subscription is a flat fee for as much everyday chatting as you want inside the app. The API is metered, so a light month is nearly free and a heavy month can dwarf any subscription. For everyday interactive use, the subscription almost always wins, and it is not close. As a rough gauge, the $20 Plus fee buys about four million GPT-5.4 Mini output tokens on the API. A typical chat exchange runs a few thousand tokens, so you could hold hundreds of conversations a month and still not spend $20 through the API. And you would lose the app, Deep Research, Sora, and everything else the subscription bundles in. The subscription is the deal precisely because OpenAI is not charging you per token for it. That gap is larger than it looks, and it points at something worth understanding before you assume a flat plan will last. An [analysis by the research firm SemiAnalysis](https://letsdatascience.com/news/chatgpt-200-plan-implies-up-to-14000-api-costs-48c1a8ca), which bought every tier and ran them to their limits, estimated that a $20 Plus plan used hard is worth roughly $700 in API-equivalent compute, and a $200 Pro plan pushed to its ceiling with agentic coding could cost OpenAI up to about $14,000 a month. In other words, the subscriptions are heavily subsidized, and the people who benefit most are the heaviest users. The API only wins when you are not chatting. If you are building a product, processing documents in bulk, or running automated agents at scale, per-token pricing (especially with Batch and cached-input discounts) is the right meter, and there is no subscription that covers programmatic use anyway. The rule of thumb: if a human is typing, buy a subscription; if code is calling the model, use the API. ## Which ChatGPT Plan Should You Choose? Strip away the tiers and the decision comes down to how hard you use the tool and whether other people share it. If you use ChatGPT occasionally, stay on **Free** and either accept the ads or switch on its ad-free mode. If you are a heavy but casual user in a price-sensitive market, **Go** is defensible, but for most people it is the trap Plus quietly solves. If you use ChatGPT for real work most days, **Plus at $20** is the answer, and it will stay the answer until you are hitting its limits daily. Only then does **Pro at $100** make sense, and only the small group who exhaust even that, heavy builders and researchers, need **Pro at $200**. The moment more than one person needs shared context, admin controls, or client-data protection, you are on **Business**, and once compliance and data residency drive the conversation you are talking to sales about **Enterprise**. In our experience helping teams put these tools to work, the plan people regret is almost never Pro. It is Go, bought to save money, that ends up delivering the least, and Plus subscriptions where nobody ever touches Deep Research, the feature that justifies the price. Before you upgrade, use what you already pay for. Before you downgrade, check whether one unused feature was the whole reason to be there. The quickest gut check is three questions. Are you hitting Plus limits before the end of most workdays? Do you regularly work with documents long enough that context length matters? Are you sharing an account with someone who should have their own? A yes to the first two points at Pro. A yes to the third points at Business. If it is no across the board, you are already on the right plan. ## How ChatGPT Pricing Compares to Claude, Gemini, and the Rest ChatGPT does not set its prices in a vacuum, and the $20 individual tier is where the whole market has clustered. ChatGPT Plus, Anthropic's Claude Pro, and Google's mid Gemini plan all sit within a dollar of each other, all aimed at the same professional user. At the top, ChatGPT Pro at $200 lines up against [Claude's $200 Max tier](https://geotoolbox.ai/blog/claude-pricing), while Google's most expensive consumer plan pushes higher still. Where ChatGPT gets undercut is at the bottom: some rivals have tested cheaper entry plans, so Go's value really depends on whether the ads and missing tools matter to you. If price is what is pushing you to look elsewhere, our roundup of the [best ChatGPT alternatives](https://geotoolbox.ai/blog/chatgpt-alternatives) sorts the cheaper options by what each one is actually good at. What that means in practice is that price is rarely the reason to pick one over another at the same tier, because they match. The difference is what each does best, ChatGPT's breadth of built-in tools, Claude's writing and reasoning, Gemini's tie-in with Google's apps. For the current numbers on the two closest rivals, the [Gemini pricing](https://geotoolbox.ai/blog/gemini-pricing) and [Grok pricing](https://geotoolbox.ai/blog/grok-pricing) guides keep the exact figures current, since those move as often as ChatGPT's do. ## Will ChatGPT Pricing Change? And Why Getting Cited Matters Yes, and OpenAI has said so. Nick Turley, its head of ChatGPT, has stated publicly that pricing will "significantly evolve" as the technology changes, and has questioned whether unlimited plans make long-term sense at all. Internal projections [reported by TechCrunch](https://techcrunch.com/2024/09/27/openai-might-raise-the-price-of-chatgpt-to-22-by-2025-44-by-2029/) back in 2024 had OpenAI eyeing $22 by 2025 and $44 by 2029. That first mark slipped, since Plus is still $20 in mid-2026, which if anything reinforces the read from the subsidy math above: today's flat fee looks less like a settled price and more like a subsidized phase. The appearance of the $100 tier and the GPT-5.6 model split are early signs of a stack being reshuffled. There is a second-order point here that matters if you publish anything. Notice how you probably arrived at a price in the first place. For a query like "chatgpt pricing," the AI answer at the top of Google is stitched together from a handful of blogs and videos, and the pages it pulls from are the ones that are current, cleanly structured, and easy to lift a number out of. The plan you pay for is also the engine deciding whether your business shows up when a buyer asks it a question. That is the part of AI pricing most people miss: you are both a customer of these models and, if you have a website, something they cite or ignore. Our guide on [how ChatGPT cites sources](https://geotoolbox.ai/blog/chatgpt-citations) covers how that selection actually works, and it is the whole reason [generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization) exists as a discipline. ## Frequently Asked Questions ### Is ChatGPT Plus worth $20 a month? For anyone who uses ChatGPT for real work most days, yes. A quick way to be sure: run a normal week and notice whether the message limits or the 10 monthly Deep Research runs ever get in your way. If they do, $20 is easily worth it; if they never do, Free or Go likely covers you. ### What's the difference between ChatGPT Free and Plus? Free runs GPT-5.6 Luna as its default model and, from the week of August 10, 2026, is rolling out uncapped text chats with a Think button, but it keeps real limits on files and images, has no access to GPT-5.6 Sol, has no full Deep Research, and shows ads in the markets where OpenAI has launched them (the EEA and Switzerland are next, confirmed for later in August 2026), though you can switch Free to an ad-free mode without upgrading, at the cost of fewer daily messages and image generation. Your chats may also be used for training unless you opt out. Plus removes the ads, moves you up to Sol, and adds the full tool suite: Deep Research, Sora, Codex, Agent Mode, and Projects. The jump is a different product, not just a bigger version of the same one. There is no legitimate way to get Plus itself for free, and there is no broad individual student discount as of mid-2026: the two-month US and Canada promotion ended on May 31, 2025 and has not returned. One narrow exception survives, a Student Plus referral programme for students at eligible universities in Australia and Colombia, which grants a single free month of Plus to new Plus sign-ups verified through a school email. Everyone else relies on Free or Go, or on institutional ChatGPT Edu access arranged by their school. ### Should I get the $100 or $200 ChatGPT Pro tier? Start with the $100 tier if you consistently hit Plus limits but do not need the absolute ceiling. It gives roughly five times the Plus quotas and heavier Codex access. The $200 tier only pays off if you also need the largest context window OpenAI gives an individual (128K on Instant, 400K on reasoning), full Sora, and roughly 20 times the Plus usage limits, which is a genuinely small group of heavy builders and researchers. ### Can I use ChatGPT Business as a solo user? No. Business requires a minimum of two seats, so the real entry cost is about $40 a month on annual billing. A solo user who wants the data-protection and admin features should look at a Pro tier or, for training exclusion specifically, weigh whether the Business two-seat minimum is worth it. ### Does ChatGPT have an annual discount or refunds? Business is where the annual discount lives: $20 per user billed yearly versus $25 monthly. The individual plans (Go, Plus, Pro) are generally billed monthly without a separate annual rate. Refund terms are limited and vary by region, so cancel before your renewal date rather than counting on a refund after it. ### Is the ChatGPT API cheaper than a subscription? For interactive chatting, no, the subscription is far cheaper because it is a flat fee for heavy use. The API is cheaper only when you are not using the chat app at all: building a product, batch-processing, or running agents, where per-token pricing with Batch and cached-input discounts is the right meter. ## The Bottom Line Most people are done at $20. Plus covers real daily work, it has held its price for three years, and the tiers above it exist for a smaller group than the marketing implies. The two moves that save the most money are the boring ones: skip Go unless the budget is genuinely tight, and actually use the plan you are on before paying for the next one up. There is one more question the price tag does not answer, though, and it is the one that matters if you run a business. Knowing what you pay to use ChatGPT is half the picture. The other half is whether ChatGPT points buyers to you when they ask it what to buy. That is a different kind of visibility, and it is the problem we built [geotoolbox's citation tools](https://geotoolbox.ai/features/citation-interceptor) to solve: seeing where AI engines mention you, where they mention a competitor instead, and what to do about it. --- ## Claude Pricing in 2026: Plans, API Costs, and Is It Worth It? > Every Claude price for August 2026: Free, Pro, Max, Team, Enterprise, and API costs for Opus 5, Sonnet 5, Haiku 4.5, and Fable 5, plus which plan is worth it. - Canonical: https://geotoolbox.ai/blog/claude-pricing - Published: 2026-07-02 · Updated: 2026-08-14 Claude costs nothing to start and up to $200 a month at the top of the individual plans, and that is before you touch the API. The consumer tiers run Free at $0, Pro at $20, and two Max tiers at $100 and $200. Teams pay per seat: Standard is about $20 to $25 per user, Premium is $100 to $125, and Enterprise runs on a seat fee plus usage. Developers pay the Claude API by the token, on a completely separate bill. Claude pricing has moved fast in 2026. Anthropic doubled Claude Code's usage limits in May, announced then paused a billing change for automated usage in June, and shipped Claude Sonnet 5 at the end of June. Most pricing guides you will find still quote a model lineup that has already changed. Below is every current price, reconciled from each vendor's own pricing page, plus the question the numbers exist to answer: which plan, if any, you should actually pay for. One thing to get straight before anything else: a Claude subscription and the Claude API are separate products with separate billing. Paying for one does not give you the other. All prices here are US list prices. ## How Much Does Claude Cost? Every Plan at a Glance Here is the whole lineup in one place, at [US prices from Anthropic](https://claude.com/pricing) as of July 2026.
PlanPrice (US)What it isBest for
Free$0/moFull chat, code generation, web search, tight usage limitsCasual, occasional use
Pro$20/mo ($17/mo annual)~5x Free usage, adds Claude Code and CoworkMost working professionals
Max 5x$100/mo5x Pro usage, priority accessPower users who hit Pro limits
Max 20x$200/mo20x Pro usage, highest limitsHeavy daily coders and researchers
Team Standard$25/seat (annual $20)Full features incl. Claude Code, 1.25x Pro usageTeams of 2 to 150
Team Premium$125/seat (annual $100)Same features + Fable 5 bundled, 5x Standard usageHeavy-usage team members
Enterprise$20/seat + usageUsage billed at API rates, compliance controlsLarge or regulated organizations
APIPay per tokenProgrammatic access to the models, no chat appDevelopers and products
Two things cause most of the confusion. The Max tier is sold as two plans, $100 and $200, that share the same models and differ only in how much you can use them. And the API in that last row is a separate product priced per million tokens, not a monthly subscription. Pro and the Team plans also bundle Anthropic's newer surfaces, Claude Code for the terminal and Claude Cowork for longer agent tasks, which matter more than the headline price once you start using them. If you are new to the platform, our guide to [what Claude is](https://geotoolbox.ai/blog/what-is-claude-ai) covers the models and tools in plain terms.
![Claude's plans in August 2026, cheapest to priciest: Free $0, Pro $20, Max 5x $100, Max 20x $200, Team Standard and Premium seats from $25 to $125 per seat for 2 to 150 people, Enterprise from $20 per seat plus usage, and a pay-per-token API covering Opus 5, Sonnet 5, Haiku 4.5 and Fable 5.](/blog/claude-pricing/claude-plans-july-2026.png)
Claude's subscription tiers, August 2026, alongside the separate pay-per-token API.
## Is Claude Free? What the $0 Plan Gives You Yes, and the free tier is a real product, not a locked demo. You get Claude on the web, the desktop apps, and mobile, with code generation, web search, memory across conversations, file creation, and extended thinking for harder problems. No credit card required. The catch is the usage limit. Free accounts get a small pool of messages that resets on a rolling five-hour window, and the ceiling shows up fast on long or heavy sessions. Anthropic does not publish an exact message count, and it shifts with load, but the pattern is that a serious work session will hit the wall within an hour or two. One feature that is not on Free is Claude Code, the terminal coding tool, which starts on Pro. For occasional use, drafting, quick questions, or testing what the model can do, Free is genuinely enough. The moment you find yourself waiting out the limit mid-task more than once a day is the signal you have outgrown it. As one long-time reviewer put it, the honest question is not whether Claude is good, it is whether you hit the free wall often enough that removing it is worth about [66 cents a day](https://jess-writes-about-tech.medium.com/is-claude-pro-worth-it-in-2026-ive-paid-for-it-for-two-years-honest-review-6407b9bc33f5). ## Claude Pro ($20/Month): The Default for Most People Pro is the plan most people should start with. It runs $20 a month, or $17 a month if you pay annually, which bills as $200 up front. That has held at $20 for a long stretch while the feature set kept growing, which makes it one of the steadier deals in AI subscriptions. What you get for it: roughly five times the usage of Free, plus the tools that turn Claude from a chat window into a work tool. Pro adds Claude Code in the terminal, Claude Cowork for longer agent-style tasks, unlimited projects to organize your chats and files, Research for multi-step reports, Claude Design, Claude Science, access to more Claude models, and the Microsoft 365 integration. It is the first tier where Claude stops feeling capped. Is $20 worth it? The useful test is not a general yes or no, it is your own week. Use Pro normally for seven days and count how often you actually hit a limit. If you rarely do, Pro is right and you do not need to spend more. If you hit the wall most days, that is real evidence for a Max tier rather than a hunch. At the same $20, Pro competes head to head with ChatGPT Plus and Gemini's mid plan, so if you are weighing them, our [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) breakdown covers where each one pulls ahead. ## Claude Max: 5x ($100) vs 20x ($200) Max is where people overthink the decision, so it helps to know what you are actually buying. Both Max tiers give you the same features as Pro, with one July 2026 exception, Fable 5, which Max bundles into its usage pool while Pro only meters it through a credit (covered below). Beyond that, the only difference is how much you can use them. Max 5x, at $100 a month, gives roughly five times Pro's usage. Max 20x, at $200 a month, gives roughly twenty times. Neither one is a smarter Claude. It is simply a bigger bucket. That reframes the whole choice. You do not upgrade to Max for a feature, you upgrade because you keep running out of Pro. Both tiers add higher output limits, earlier access to new features, and priority access when Claude is under heavy load, but usage capacity is the real product. The practical advice from people who have paid for all of these is the same every time: start on Pro, and only move up when you can feel the limits getting in your way. Then take the smallest step that fixes it. Most people who hit Pro's ceiling are well served by Max 5x, and only heavy daily users, the kind running long coding sessions or agent workflows, need the 20x tier. There is one wrinkle worth knowing: because usage is per account, some heavy users find two separate Max 5x accounts give them more parallel room than a single Max 20x, though that means juggling two logins. For most people it is not worth the hassle. Claude's usage limits have moved more than once in 2026, and Anthropic does not publish exact message counts. Treat any specific "X prompts per window" figure you see as an estimate, and check the current allowance on your plan before assuming a tier will cover your volume. ## Claude Code Pricing: Is the $20 Tier Worth It? Claude Code, Anthropic's terminal coding agent, is the single biggest reason people search for Claude pricing, so it deserves a clear answer. The short version: Claude Code is included in Pro at $20, it was not removed, and there is no separate Claude Code subscription. What trips people up is that Claude Code draws from the same usage pool as everything else on your plan, so if you keep hitting the wall, [cutting Claude Code's token usage](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) can matter as much as the tier you pay for.
PlanIncludes Claude Code?Usage vs ProBest for
FreeNo-Not for coding
Pro ($20)Yes1x (baseline)Light coding, small repos
Max 5x ($100)Yes5xRegular projects
Max 20x ($200)Yes20xDaily heavy coding, agent runs
Team Standard ($25/seat)Yes1.25x ProMixed teams
Team Premium ($125/seat)Yes5x StandardHeavy-usage seats
So why does Claude Code feel expensive? Because coding burns tokens far faster than chatting. A single agent task can read your whole codebase, run tools, and think through several attempts, and all of that comes out of the same five-hour and weekly budget your chats use. On Pro, a real coding session can hit the wall quickly. That is the specific pressure that pushes developers to Max 20x, where the larger bucket is built for daily heavy use. If your usage is spiky or you are building software rather than coding interactively, there is a third path: skip the subscription and use the API directly, billed per [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai). One heuristic from teams tracking this closely is that a Max 20x plan tends to beat raw API billing once you pass [roughly 70 million tokens a month](https://www.finout.io/blog/claude-code-pricing-2026) of typical Sonnet-heavy coding, and the API wins below that or when your usage is uneven. The threshold moves materially once you shift to Opus. One warning that catches people out: if you have `ANTHROPIC_API_KEY` set in your shell, Claude Code bills at API rates and ignores your subscription entirely. ## Claude Team and Enterprise: Per-Seat Pricing Once you are buying for a group, the plans change shape. Claude Team is built for organizations of 2 to 150 people and requires a minimum of two seats, so the practical entry point is twice the seat price. It comes in two seat types you can mix on the same team.
Team StandardTeam PremiumEnterprise
Price$25/seat (annual $20)$125/seat (annual $100)$20/seat + usage
Usage per seat1.25x Pro6.25x Pro (5x Standard)Negotiated
Claude Code and CoworkYesYesYes
SSO and admin controlsYesYesYes
SCIM, audit logs, complianceNoNoYes
The key thing to understand is that the seat types are nearly the same product; the main difference is usage. Standard and Premium seats both include Claude Code, Cowork, and every model, with Fable 5 the one exception since July 2026: Premium bundles it into the usage pool, while Standard (like Pro) provides a one-time credit and then meters it at API rates. What separates them is how much you get: a Standard seat is about 1.25 times a Pro plan's per-session usage, while a Premium seat is 6.25 times. So the real decision is not who gets the coding tool, it is who needs the headroom. Put light users on Standard, put the people who run Claude Code all day on Premium, and mix the two on one bill. All Team plans keep your content out of model training by default, which is often the real reason a company moves off individual accounts. One thing to know about the limits: they are per seat, not pooled, so a heavy user cannot borrow a colleague's unused capacity. Enterprise is the tier for larger or regulated organizations. The pricing page lists it as a seat price plus usage billed at API rates, with the real terms set through sales. What you are paying for is the governance layer Team does not include: role-based access control, SCIM provisioning, audit logs, a compliance API, custom data retention, IP allowlisting, a HIPAA-ready option, and Claude Security. If regulators rather than budgets are driving the conversation, that is Enterprise territory. ## Claude API Pricing: Pay per Token (Sonnet 5 and Fable 5 Included) The API is a different product with a different meter. There is no monthly fee and no message cap. You pay per token, split into input (what you send) and output (what the model returns), priced per million tokens. This is where Anthropic's lineup moves fastest, and where most pricing guides are already out of date: Claude Sonnet 5 launched at the end of June 2026, and few guides have it yet. Here are the [current published rates](https://platform.claude.com/docs/en/about-claude/pricing) for the models Anthropic recommends today.
ModelInput (per 1M)Output (per 1M)Best for
Haiku 4.5$1$5High-volume, simple tasks
Sonnet 5$2$10The production default
Opus 5$5$25Complex reasoning, agentic coding
Opus 4.8 (legacy)$5$25Same rate; superseded by Opus 5
Fable 5$10$50Long-running autonomous agents
A few things to read into that table. [Claude Sonnet 5](https://geotoolbox.ai/blog/claude-sonnet-5) runs $2 input and $10 output per million tokens; Anthropic announced that as introductory pricing through August 31, 2026, then made it permanent in August 2026, cancelling the planned rise to $3 and $15. It is the sensible default for most production work. The Opus tier is now led by [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5), released July 24, 2026 at the same $5 and $25 rate as Opus 4.8, which remains available as a legacy model. Fable 5, at $10 and $50, is Anthropic's most capable model, built for long-running autonomous agents, which is why it costs double Opus and is overkill for a job that Opus or Sonnet can handle. There is also a limited-availability Claude Mythos 5 at the same rate. Unless you are running long unattended agent workloads, Sonnet 5 or Opus 5 is almost always the right call. One subscription note, because it changed on July 20, 2026: Fable 5 is not included across the consumer plans the way Opus 5 and Sonnet 5 are. It comes bundled on Max and Team Premium at up to 50% of your weekly usage limits, while Pro and Team Standard get a one-time $100 usage credit and then pay these API rates ($10 / $50) to keep using it. So on Pro, Fable is a metered add-on rather than an included model. The full story, including why it was briefly pulled and brought back, is in our [Claude Fable 5 explainer](https://geotoolbox.ai/blog/fable-5-ban). Two nuances hide real money here. First, cheaper per token is not the same as cheaper per task. Sonnet 5's rate sits well below Opus 5, so for straightforward work it is the cheaper choice, but it tends to use more reasoning and tool calls on agent workloads, which narrows the gap. On a heavy agent job that burns through enough extra tokens and tool calls, Sonnet 5 can end up costing close to the same job on Opus. Separately, the newer models (Sonnet 5, Opus 4.8, Fable 5) count tokens with an updated tokenizer that runs about 30% higher than older models like Sonnet 4.6, so a team moving up from an older Sonnet should budget against its actual token volume, not the per-token rate alone. Two discounts then cut the bill hard: the Batch API takes 50% off both input and output for work that can run asynchronously, and prompt caching drops repeated context to about 10% of the input rate, savings of up to 90%. Server-side tools bill on top, for example web search at $10 per 1,000 queries, and a fast mode for Opus 5 and Opus 4.8 runs at $10 input and $50 output per million, double the standard rate, for lower latency (a research preview, on the Anthropic API only). ## Subscription vs API: Which Is Cheaper? The two meters are built for different jobs. A subscription is a flat fee for as much interactive use as your limits allow. The API is metered, so a quiet month is nearly free and a busy one can dwarf any subscription. The rule of thumb is simple: if a human is typing, buy a subscription; if code is calling the model, use the API. For everyday chatting and coding the subscription wins easily, because Anthropic is not charging you per token for it. For a product that processes documents in bulk or runs agents at scale, per-token pricing with batch and caching discounts is the right meter, and no subscription covers programmatic use at scale anyway. That last point is worth a careful note, because it nearly changed in 2026. Anthropic [announced](https://support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan) that starting June 15, programmatic usage through the Agent SDK, `claude -p`, and third-party apps would move out of the shared subscription pool onto a separate monthly credit billed at API rates. Then it paused the change before it took effect. As of now, those surfaces still draw from your Pro, Max, or Team limits exactly as before, there is no separate credit, and Anthropic has said it will give advance notice before any future version. If you build automated workflows on a subscription, this is the line to watch. ## The Real Cost Is the Usage Limits The prices above are the easy part. What actually determines whether a plan works for you is the usage limits, and they are less visible than the sticker price. Claude enforces two at once: a five-hour rolling window and a weekly cap. The weekly one is not published as a hard number, and there is no live meter telling you how much you have left, so the ceiling tends to arrive as a surprise mid-task. And because that budget is shared across chat, Claude Code, and Cowork, a heavy afternoon of coding can lock you out of ordinary chat for the rest of the window. There is also a quieter limit inside the limit, and it does not work the way it used to. Opus no longer carries its own separate cap: Anthropic removed the Opus-specific weekly limit, so Opus now draws on the overall allowance. It is Sonnet that has the extra ceiling — Max plans run "two weekly usage limits: one that applies across all models and another for Sonnet models only." If you read older guides describing a tight Opus-only quota, they predate that change. The good news is the direction of travel. In May 2026, Anthropic [doubled Claude Code's five-hour limits](https://www.anthropic.com/news/higher-limits-spacex) across Pro, Max, and Team, removed the peak-hours penalty, and raised API rate limits for Opus, on the back of a large compute deal with SpaceX. In our experience helping teams adopt these tools, the plan people regret is almost never the expensive one. It is the cheap tier bought to save money that ends up capping the work, and Pro subscriptions where nobody ever touches the features that justify the price. ## Is Claude Worth It? Claude vs ChatGPT, Gemini, and Grok on Price On the consumer side, price is rarely the deciding factor, because the whole market has clustered at the same number. Claude Pro, ChatGPT Plus, and Google's mid Gemini plan all sit within a dollar of each other at around $20, and Claude's $200 Max tier lines up against ChatGPT's $200 Pro tier. The one that prices differently is Grok, whose main tier sits higher at $30.
ProviderMain paid planFlagship API (input / output per 1M)
Anthropic ClaudeClaude Pro $20Opus 5, $5 / $25
OpenAI ChatGPTChatGPT Plus $20GPT-5.6 Sol, about $5 / $30
Google GeminiGoogle AI Pro about $20Gemini 3.1 Pro, about $2 / $12
xAI GrokSuperGrok $30grok-4.6, about $2 / $6
Because the headline prices match, the choice comes down to what each does best rather than what it costs: Claude's writing and reasoning, ChatGPT's breadth of built-in tools, Gemini's tie-in with Google's apps, Grok's real-time access to X. On the API, Claude is not the cheapest on raw list price, but its prompt caching is more aggressive than most rivals', which changes the real cost for apps that reuse context. Gemini's Flash models undercut it on raw tokens, though the real bill depends on the same hidden meters, as our [Gemini API pricing](https://geotoolbox.ai/blog/gemini-api-pricing) guide breaks down. For the current numbers on the closest rivals, our [ChatGPT pricing](https://geotoolbox.ai/blog/chatgpt-pricing), [Gemini pricing](https://geotoolbox.ai/blog/gemini-pricing), and [Grok pricing](https://geotoolbox.ai/blog/grok-pricing) guides keep the exact figures current, since those move as often as Claude's do. ## Which Claude Plan Should You Actually Pay For? Most people overbuy. Here is how to decide, from the bottom up. **Stay free** if you use Claude occasionally and do not code with it. The free tier covers it, and you will know you have outgrown it the day you keep hitting the limit mid-task. **Pay $20 for Pro** if you use Claude for real work most days. This is the right answer for the large majority of paying users, and it stays the answer until you are hitting its limits daily. **Pay $100 for Max 5x** once you consistently run out of Pro, and step up to **$200 for Max 20x** only if you exhaust even that, which really means daily heavy coding or agent runs. **Buy Team** the moment more than one person needs shared billing, admin controls, or data kept out of training, and give Premium seats to whoever needs the most usage. **Talk to sales about Enterprise** when compliance and data residency, not budget, drive the decision. **Use the API instead of a plan** if you are building software. And if you are a student, check whether your school has an institutional Education plan before paying out of pocket. ## Why Claude's Price Matters for Your Brand One point here outlasts any specific price, and it is the one that matters if you run a business. Whichever tier people pay for, Claude is not just a tool they visit. It answers questions for millions of people, and increasingly those questions are about what to buy and who to hire. When someone asks Claude to recommend a product in your category, it names some companies and leaves out others. So the more useful question is not which plan you should buy. It is whether Claude mentions your brand at all when a buyer asks it for a recommendation, and what it says when it does. That visibility does not come on a pricing tier, and it is the gap we help businesses close. Our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) maps where Claude and the other AI engines cite sources your brand is missing from, so you can see which conversations to get into, and you can learn the broader method in our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). Knowing what you pay to use Claude is half the picture. Whether Claude points buyers to you is the other half, and it is what [generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization) exists to solve. ## Frequently Asked Questions ### How much does Claude cost per month? Claude is free to start. Paid individual plans are Pro at $20 a month (or $17 billed annually), Max 5x at $100, and Max 20x at $200, as of July 2026. Team seats run $25 for Standard and $125 for Premium, cheaper on annual billing, and the developer API is billed separately per token. ### Is Claude Pro worth $20 a month? For anyone who uses Claude for real work most days, yes. A quick test: use it normally for a week and notice whether the usage limits ever get in your way. If they do, $20 is easily worth it; if they never do, the free tier probably covers you. Pro also adds Claude Code and Cowork, which Free does not have. ### What's the difference between Max 5x and Max 20x? Mostly how much you can use them. Both give the same features as Pro and, unlike Pro, bundle Fable 5 into the usage pool rather than metering it after a one-time credit; Max 5x ($100) is about five times Pro's usage and Max 20x ($200) is about twenty times. Neither is a smarter Claude. Start with 5x if you are hitting Pro's limits and move to 20x only if you exhaust that too. ### Is the $20 Claude Code worth it, and did it get removed? Claude Code is still included in Pro at $20; it was not removed, and there is no separate Claude Code plan. It feels expensive because coding burns your shared usage budget fast, which is what pushes heavy users to Max 20x or to the API. For occasional coding on small projects, the $20 Pro tier is genuinely enough. ### Is Claude cheaper than ChatGPT? At the main tier they match: Claude Pro and ChatGPT Plus are both about $20. On the API it depends on the model. Claude Opus 5 runs about $5 and $25 per million tokens, close to [GPT-5.6](https://geotoolbox.ai/blog/gpt-5-6), while Claude Sonnet 5 at $2 and $10 undercuts it for most production work. ### How do Claude's usage limits actually work? Two limits apply at once: a five-hour rolling window and a weekly cap. Anthropic does not publish exact numbers, there is no live meter, and all your usage (chat, Claude Code, Cowork) shares one pool. Opus no longer has its own cap and draws on the overall allowance; Sonnet is the model with a separate weekly limit on Max plans. Switching to a lighter model is a manual step (`/model`), not an automatic fallback. Anthropic doubled Claude Code's five-hour limits in May 2026. ## Sources - Plans & Pricing - Claude by Anthropic - `claude.com/pricing` - Choose a Claude plan - Claude Help Center - `support.claude.com/en/articles/11049762-choose-a-claude-plan` - Pricing - Claude Platform Docs - `platform.claude.com/docs/en/about-claude/pricing` - Introducing Claude Sonnet 5 - Anthropic - `anthropic.com/news/claude-sonnet-5` - Higher usage limits for Claude and a compute deal with SpaceX - Anthropic - `anthropic.com/news/higher-limits-spacex` - Use the Claude Agent SDK with your Claude plan - Claude Help Center - `support.claude.com/en/articles/15036540-use-the-claude-agent-sdk-with-your-claude-plan` - Claude Code Pricing 2026 - Finout - `finout.io/blog/claude-code-pricing-2026` --- ## Gemini API Pricing in 2026: Every Model, Tier, and Hidden Cost > Gemini API pricing 2026: every model's token cost, the four service tiers, the free tier, Vertex vs Developer API, and the hidden costs that inflate your bill. - Canonical: https://geotoolbox.ai/blog/gemini-api-pricing - Published: 2026-07-02 · Updated: 2026-08-14 Gemini API pricing starts at about $0.10 per million tokens with no subscription, which makes it look like one of the cheapest ways to build on a frontier model. It often is. But the per-token rate is not the bill. The same model sells at four different service-tier prices, the free tier comes with a real catch, and a handful of meters (thinking tokens, a context cliff, and grounding) decide what you actually pay. This is the developer's cost reference for the Gemini API: every model's token price, the service tiers, the free-tier limits, Vertex AI versus the Developer API, and the hidden costs that inflate the number. If you want a monthly consumer plan instead of per-token billing, see the [consumer Gemini pricing](https://geotoolbox.ai/blog/gemini-pricing) breakdown instead. ## How Much Does the Gemini API Cost? Every Model's Token Price The Gemini API charges you [per token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), the chunks of text a model reads and writes, billed separately for input tokens and output tokens. There is no subscription. You pay for what your software sends and receives, and the rate depends entirely on which model you call. Here are the current pay-as-you-go rates for the models worth using, at Standard-tier prices as of July 2026.
ModelInput (per 1M tokens)Output (per 1M tokens)Best for
Gemini 2.5 Flash-Lite$0.10$0.40Cheapest option, high-volume classification
Gemini 3.1 Flash-Lite$0.25$1.50Prior-generation budget tier
Gemini 2.5 Flash$0.30$2.50Multimodal workhorse
Gemini 3.5 Flash-Lite$0.30$2.50New budget tier, fastest in the 3.5 line (July 2026)
Gemini 3.6 Flash$1.50$7.50New default workhorse, successor to 3.5 Flash (July 2026)
Gemini 3.5 Flash$1.50$9.00Superseded by 3.6 Flash
Gemini 2.5 Pro$1.25 / $2.50$10.00 / $15.00Higher rate over 200K-token prompts
Gemini 3.1 Pro (Preview)$2.00 / $4.00$12.00 / $18.00Top Pro model, paid only, over-200K premium
Three patterns run through that table. Flash-Lite is the floor and Pro is the ceiling, and the gap is wide: Gemini 3.1 Pro costs eight times more per input token than 3.1 Flash-Lite. For classification, extraction, and routine summarizing, the cheap models are usually enough, and routing simple work to Flash-Lite is where most teams find their savings. Output costs more than input, four to more than eight times on the Flash models, so long responses are where bills grow fastest. The Pro models also carry a context cliff: once a single prompt crosses 200,000 tokens, the input rate roughly doubles. More on that below. One naming note, since older guides get it wrong. As of July 21, 2026, the current default workhorse is **Gemini 3.6 Flash** at $1.50 input and $7.50 output, the successor to 3.5 Flash (which stays live at the higher $9.00 output but is now superseded). Google also shipped **Gemini 3.5 Flash-Lite** at $0.30 / $2.50 as a newer budget tier. The rates come straight from [Google's Gemini API pricing page](https://ai.google.dev/gemini-api/docs/pricing), the source of record, which changes often enough to check before you commit a budget. For how the two new models compare head to head, see our [Gemini 3.6 Flash vs 3.5 Flash-Lite](https://geotoolbox.ai/blog/gemini-3-6-flash-vs-3-5-flash-lite) guide. ## Standard, Batch, Flex, and Priority: The Four Service Tiers The headline rate is only one of four prices for the same model. Since April 2026, every Gemini model is billed on a service tier, and the tier you pick swings the bill from half price to nearly double. Most pricing guides still show only Standard and Batch, which is why a number you read somewhere may not match what you are charged. Here is Gemini 3.6 Flash across all four tiers.
TierInput (per 1M)Output (per 1M)What you trade
Standard$1.50$7.50Baseline price, normal latency
Batch$0.75$3.7550% off, up to 24 hours to return
Flex$0.75$3.75About 50% off, latency-tolerant, can be queued
Priority$2.70$13.50Roughly 1.8x Standard, non-sheddable traffic
Standard is the default. Batch runs any model at half price if you submit work as a job and can wait up to a day for it, which suits overnight enrichment and bulk generation. Flex is the middle ground: close to Batch pricing for on-demand calls you are willing to let the system queue during busy periods. Priority is the premium lane, about 1.8 times Standard, for production traffic that needs predictable latency: Google doesn't guarantee a specific response time, but Priority requests are non-sheddable and won't get queued behind Standard traffic during busy periods, the way Flex can. The practical trap is reading a number without knowing its tier. If an AI Overview tells you Gemini 3.6 Flash costs $2.70 per million input tokens, it is quoting the Priority rate, not the $1.50 headline. Always confirm which tier a quoted price belongs to before you budget against it. ## The Gemini API Free Tier: What's Free, and the Catch Yes, the Gemini API has a free tier, and no, it is not a trial that expires. Through Google AI Studio you can call the Flash and Flash-Lite models with no credit card. Google no longer publishes a fixed public table of the limits; the Flash-tier allowance runs into the low thousands of requests a day, but the number that matters is the live quota AI Studio shows for your project. Either way it is plenty for prototyping and light production. Two limits decide whether the free tier fits your project. The first is scope. Since April 2026 the Pro models have effectively left the free tier, and 3.1 Pro in particular is paid from the first call. If your workload runs on Flash and Flash-Lite, the free tier stretches a long way. If it needs Pro-grade reasoning, budget for paid from day one. The second is the data trade, and it is the one that surprises teams doing client work. On the free tier, [Google's API terms](https://ai.google.dev/gemini-api/terms) allow it to use your inputs and outputs to improve its models, and human reviewers may read them. That is fine for a weekend prototype and wrong for anything confidential. Paid usage, on either the Developer API or Vertex AI, is not used for training. If you are sending customer data, treat the free tier as off-limits. One more thing that surprises people: extra API keys do not add quota. Rate limits are enforced per project, not per key, so spinning up a second key in the same project buys you nothing. To raise limits you move up a usage tier, which is the next section. ## Rate Limits and Usage Tiers: Free to Tier 1, 2, and 3 Your rate limits are not fixed. They rise as your project climbs a usage-tier ladder tied to how much you have spent, and each tier also carries a billing cap that pauses service if you hit it.
TierHow you reach itBilling cap
FreeActive project or free trialN/A
Tier 1Link a billing account$250
Tier 2$100 spent, plus 3 days since first payment$2,000
Tier 3$1,000 spent, plus 30 days$20,000 to $100,000+
Higher tiers raise your requests-per-minute, tokens-per-minute, and requests-per-day limits. Google does not publish those numbers as a single public table; you view your active limits inside AI Studio, and they lift automatically as you cross each spending threshold. The [billing documentation](https://ai.google.dev/gemini-api/docs/billing) spells out the caps: when your cumulative spend hits a tier limit, service pauses for every project on that billing account until the next cycle. This is also where the dreaded 429 lives. A `RESOURCE_EXHAUSTED` error means you have hit a rate limit, and it can strike even on a paid key when a project has not finished provisioning its higher quota, or when an image model is still pinned to free-tier limits. The fix is usually to confirm billing is fully linked, give the project time to provision, and add backoff-and-retry rather than hammering the endpoint. ## The Costs That Surprise You: Thinking Tokens, the Context Cliff, and Grounding The sticker rate tells you what a token costs. It does not tell you how many tokens you will be billed for, and that is where real bills diverge from estimates. Four things drive the gap. First, thinking tokens are billed as output. Gemini's reasoning models generate internal thinking before they answer, and those tokens bill at the full output rate even though the user never sees them. A model can return a two-sentence answer and charge you for thousands of tokens of hidden reasoning. It is why a low-headline-rate model can cost more in practice than a pricier one that reasons less: the effective cost depends on how verbose the thinking is on your workload. Setting a thinking budget caps how many tokens the model spends reasoning before it must answer. Second, the 200K context cliff doubles Pro input. On the Pro models, once a single prompt crosses 200,000 tokens of [context](https://geotoolbox.ai/glossary/context-window), the input rate roughly doubles and output climbs with it: Gemini 3.1 Pro input goes from $2.00 to $4.00 per million, output from $12.00 to $18.00. Retrieval pipelines that stuff large documents into every call cross that line on every request without anyone noticing, because dashboards show averages and the cliff hides in the few oversized prompts. Third, grounding with Google Search is a separate line. Letting a model check live Search is not free tokens; it is a per-query charge. The [Gemini 3 family](https://ai.google.dev/gemini-api/docs/pricing) gets 5,000 grounded prompts a month free, shared across the family, then $14 per thousand. The 2.5 models get 1,500 a day free, then $35 per thousand. Grounding with Google Maps runs $25 per thousand on the 2.5 models and shares the Gemini 3 family's $14 rate. A single request can also fire more than one billable search, so the line item grows faster than the request count. Fourth, audio input costs more than text. On the same model, audio input is priced above text input, often three times higher. A voice feature is not billed like a text feature, even before you add the separate audio-output meters. ## Multimodal Pricing: Images, Video, and Audio Generating media is metered on its own scales, not in text tokens, and the [media rates](https://ai.google.dev/gemini-api/docs/pricing) look nothing like the token table. Feeding media in is metered too, as tokens. An input image costs a fixed count by size, roughly 560 tokens for a small one and over 1,100 for a large, and a PDF is billed per page as an image plus its extracted text. So a vision or document-processing app has a bigger input line than its character count suggests, and it is worth counting image and page tokens before you ship. Images out run through Google's Flash Image models. The original, nicknamed Nano Banana, is Gemini 2.5 Flash Image at about $0.039 per image. Its successor, Nano Banana 2 (Gemini 3.1 Flash Image), bills as image-output tokens at $60 per million, which works out to roughly $0.045 for a 0.5K image up to $0.151 for 4K. If you generate at volume, resolution is a real cost lever, not a detail. Video is the most expensive meter. Veo 3.1 costs $0.40 per second at 720p or 1080p and $0.60 per second at 4K on the Standard model, with cheaper Fast and Lite variants. A single ten-second 4K clip is $6 before you iterate, so video generation belongs behind a hard budget. Audio has its own meters again. The Live API and the text-to-speech models bill separately, with audio output priced well above text. Live Translate, for example, runs about $3.50 per million tokens of input and $21 per million tokens out. Embeddings are the cheap corner of the catalog. Gemini Embedding 001 is $0.15 per million tokens, and the newer Embedding 2 is $0.20, with batch pricing halving both. For a retrieval system, embedding cost is usually a rounding error next to the generation calls it feeds. If your product touches images, video, or voice, model those meters separately. They will not show up in a token estimate, and video in particular can dwarf everything else on the bill. ## Context Caching and Batch: How to Actually Cut the Bill Two levers cut a Gemini bill more than any model swap, and teams routinely miss both. Batch is the easy 50%. Any job that does not need an instant answer (overnight enrichment, bulk classification, offline generation) can go through the Batch tier for half price in exchange for up to 24 hours of latency. It is one of the easiest large savings to skip, and it needs no code beyond submitting the work as a batch job instead of a live call. Caching comes in two forms, and the difference decides whether it costs you anything. Implicit caching is automatic on Gemini 2.5 and newer models: when a request reuses a prefix you have sent before, Google passes on the discount, roughly 10% of the input rate, with no storage fee and nothing to manage. Explicit caching is the version you create and hold on purpose, which guarantees the discount for a system prompt or codebase you know you will reuse but adds a storage meter, about $1.00 per million tokens an hour on Flash and $4.50 on Pro, running whether or not you use the cache. That storage meter gives explicit caching a break-even. A cached 200K-token context on Pro costs around $0.90 an hour to hold and saves about $0.36 on each cache hit, so it pays for itself at roughly two or three reuses an hour and loses only when a large context sits nearly idle. Lean on implicit caching by default; reach for explicit caching on the big, hot contexts where you want the discount guaranteed, and let cold ones expire. Stacking the levers is where the big numbers come from. On Vertex AI, [cached input is priced at 10% of the standard input rate](https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing), a 90% reduction, alongside an advertised 50% batch discount, and a high-reuse batch job can combine both. Two more habits help: route classification and extraction to Flash-Lite instead of Pro, and set thinking budgets so reasoning tokens cannot balloon a simple task. ## A Worked Example: What a Real Gemini App Actually Costs Numbers in a table are abstract until you stack every meter into one bill. Take a support chatbot on Gemini 3.5 Flash handling 100,000 conversations a month. (The same workload on its successor, 3.6 Flash, runs the output-driven lines about 17% cheaper on the output-rate change alone, $7.50 instead of $9.00; the shape of the bill is identical.) Each conversation ships a 20,000-token system prompt and knowledge base, a short user turn, returns a 400-token answer, spends about 1,000 tokens thinking, and runs one grounded Search. Here is how the estimate and the real bill diverge.
Line itemHow it is billedMonthly cost
Input tokens~20,500 x 100K at $1.50/1M~$3,075
Visible output400 x 100K at $9.00/1M~$360
Thinking tokens1,000 x 100K at the $9.00 output rate~$900
Grounding~95K searches at $14/1,000 (after 5K free)~$1,330
Sticker estimate (input + visible output)what a naive calculator shows~$3,435
Real bill (all meters)the number that actually posts~$5,665
The sticker estimate misses by about 65%. Thinking tokens and grounding, neither of which appears in a per-token calculator, add more than $2,000 a month on their own. This bot grounds on every turn, which is the high end; one answering mostly from its own knowledge base would ground less and see a smaller gap. The shape holds either way: the meters a calculator ignores are the ones that move the bill. Now apply the levers. That 20,000-token system prompt is identical on every call, so cache it. Cached input drops to roughly 10% of the rate, cutting the input line from about $3,075 to about $375, while the storage meter for a context this heavily reused costs only around $15 a month. The real bill falls from about $5,665 to about $3,000, nearly a 50% cut, from one change. If any part of the workload were offline rather than live chat, moving it to the Batch tier would halve it again.
![Waterfall chart of a Gemini 3.5 Flash chatbot's monthly bill: a per-token calculator shows about $3,435 from input and visible output, but the real bill is about $5,665 once thinking tokens ($900) and grounding ($1,330) are added, then drops to about $3,000 with context caching.](/blog/gemini-api-pricing/anatomy-gemini-api-bill.png)
The two meters a token calculator ignores, thinking tokens and grounding, add about $2,200 a month, until caching the shared context cuts the bill nearly in half.
Model every meter before you ship, then attack the two or three lines that dominate, which are rarely the ones the token table points at. ## Gemini Developer API vs Vertex AI: Same Tokens, Different Bill You can reach the same Gemini models two ways, and the per-token rates are close to identical. What differs is everything wrapped around the tokens. The Gemini Developer API, through Google AI Studio, is the simple path. You get an API key, the free tier lives here, and you can be live in minutes. It is the right choice for most projects, prototypes, and anything that does not need enterprise controls. Vertex AI, which Google renamed the Gemini Enterprise Agent Platform in April 2026 and which most people still call Vertex, serves the identical models through Google Cloud. Its base rates match the [Developer API](https://cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing) on global endpoints, though data-residency and other non-global endpoints can carry a small regional uplift. What it adds is the machinery large deployments need: SLAs, VPC Service Controls, compliance certifications, and IAM and billing folded into the rest of your Google Cloud account. It is also where new models tend to land first, though the unreleased [Gemini 3.5 Pro](https://geotoolbox.ai/blog/gemini-3-5-pro) is not on its published pricing page either. Vertex's headline advantage for high-volume production is Provisioned Throughput, which reserves guaranteed capacity at a flat hourly rate. A committed reservation earns a double-digit percentage discount, larger on a one-year term than a one-month one, and beats pay-as-you-go once your traffic is high and steady. The trade-off is that Vertex moves you into Google Cloud's billing surface, and that surface has its own costs: regional endpoint uplift, provisioned-throughput commitments, logging and storage, and network egress add up in ways the token rate never hints at. Teams that expected "same price as the API" are the ones surprised by a Vertex bill, and it is usually the cloud plumbing, not the model, doing the damage. The rule of thumb: build on the Developer API until you need provisioned capacity, data-residency guarantees, or compliance sign-off. At that point Vertex earns its complexity. Below it, the extra surface is cost and effort you do not need. ## How Gemini API Pricing Compares to OpenAI and Anthropic On raw token price, Gemini's Flash models are among the cheapest credible options for high-volume work. Here is each provider's cheap tier against its frontier model, at Standard rates.
ProviderCheap tier (in / out per 1M)Frontier (in / out per 1M)
Google Gemini2.5 Flash, $0.30 / $2.503.1 Pro, $2.00 / $12.00
OpenAIGPT-5.4-mini, $0.75 / $4.50GPT-5.5, $5.00 / $30.00
Anthropic ClaudeHaiku 4.5, $1.00 / $5.00Opus, $5.00 / $25.00
Gemini 2.5 Flash at $0.30 input and $2.50 output is a fraction of what the frontier models charge, and the fair fight is Flash against each rival's own cheap tier: GPT-5.4-mini at [OpenAI's published rates](https://developers.openai.com/api/docs/pricing) and Claude Haiku. Even there, Gemini's $0.30 input undercuts both. You can read the full breakdowns in our [ChatGPT API pricing](https://geotoolbox.ai/blog/chatgpt-pricing) and [Claude API pricing](https://geotoolbox.ai/blog/claude-pricing) guides. There is a subtler wrinkle no table row captures: the 200K cliff erodes Gemini's long-context edge. Gemini 3.1 Pro at $2.00 input undercuts GPT-5.5 by more than half on prompts up to 200,000 tokens. Cross that line and Gemini's input doubles to $4.00 while OpenAI and Anthropic hold their rates flat, so the gap narrows sharply, though even at $4.00 Gemini 3.1 Pro still sits under GPT-5.5's $5.00. Where it can actually flip is against a cheaper-tier rival: Claude Sonnet 5 undercuts Gemini's post-cliff $4.00 at its $2.00 input rate (which Anthropic made permanent in August 2026, cancelling the planned $3.00 rise). A workload that lives in very long context can erase Gemini's short-context price win, so price the cliff into your own token distribution before you commit. And remember the reasoning-token tax applies to all three. Every provider charges for internal thinking on its reasoning tiers, so a base-rate comparison can flip once you measure how verbose each model is on your actual prompts. Benchmark two or three models on your real workload before you commit. The lowest number in a table is a starting point, not the bill. ## The Deprecation Treadmill: What's Changing in 2026 Gemini's lineup turns over fast, and pricing a project against a model that is about to disappear is a common, expensive mistake. Here is the current state of play from [Google's deprecation schedule](https://ai.google.dev/gemini-api/docs/deprecations). Gemini 2.0 Flash and 2.0 Flash-Lite shut down on June 1, 2026, so any guide still listing them as budget options is out of date. The bigger event is ahead: the entire 2.5 series (Pro, Flash, and Flash-Lite) retires on October 16, 2026, with 2.5 Pro pointing to 3.1 Pro, 2.5 Flash to 3.5 Flash (now itself superseded by 3.6 Flash), and 2.5 Flash-Lite to 3.1 Flash-Lite. On the media side, Imagen 4 shuts down on August 17, 2026, replaced by the Gemini image models. The risk is the migration path, because the wrong replacement can multiply your cost. Moving a cheap 2.5 Flash-Lite workload to 3.5 Flash instead of 3.1 Flash-Lite jumps you from $0.10 to $1.50 input, fifteen times the price, for work that never needed a frontier model. Same-class upgrades follow the tier, cheap to cheap and mid to mid, which keeps the jump small. Map every model you depend on to its named successor before its shutdown date, and confirm the replacement is the same class, not the next one up. ## Set a Hard Spend Cap Before You Get a Surprise Bill The bill-shock stories are real: a retry loop left running overnight, a leaked key mining tokens for hours, a batch job that overshoots. A Google Cloud budget alert will not save you from them, because an alert only notifies; it does not stop spend. By the time the email arrives, the money is gone. The controls that actually cap spend are more specific. The usage-tier billing caps pause service once your cumulative spend hits the tier limit, so a Tier 1 billing account cannot run past $250 in a cycle. Since March 23, 2026, Google has also rolled out prepay billing: linking a billing account requires prepaying a minimum $10 credit balance, and requests only serve while that balance stays positive, which acts as a de facto hard cap once you stop topping it up. Inside AI Studio you can also set per-project spend caps, useful when several projects share one billing account, though [Google warns](https://ai.google.dev/gemini-api/docs/billing) that batch jobs and agent sessions can overshoot a project cap slightly because billing lags real usage by about ten minutes. Beyond the platform controls, the operational habits matter more. Watch cost per call, not just total spend, because two identical-looking requests can bill very differently once thinking tokens and context length vary, and averages hide the expensive few. Put a ceiling on retries so a failing loop cannot run away. And guard your API key like a credential, because a stolen key is a direct line to your billing account. The teams that stay in control track effective cost per request from day one, rather than reading the sticker rate and hoping. ## The Sticker Price Is Where the Bill Starts The Gemini API is genuinely cheap at the low end and genuinely easy to misjudge everywhere else. The per-token table is the opening number. The service tier, the free-tier data trade, thinking tokens, the 200K cliff, grounding, and the storage meter behind caching are what turn it into an invoice. Price a project against all of them, not just the headline, and the surprises mostly disappear. One shift worth noticing: these exact prices are increasingly quoted back to developers by AI answers. When someone asks ChatGPT, Gemini, or [Google's AI Mode](https://geotoolbox.ai/blog/google-ai-mode-seo) what the Gemini API costs, an engine is choosing which page to cite. In our experience tracking which sources those engines pull from, the pricing pages that win are the clear, current, and machine-readable ones. If you publish anything developers price decisions against, it is worth knowing whether the engines cite you or a competitor, which is what [geotoolbox](https://geotoolbox.ai/tools/ai-readiness) checks. ## Frequently Asked Questions ### Is the Gemini API free? There is a free tier through Google AI Studio with no credit card, but it is in practice limited to the Flash and Flash-Lite models at roughly a thousand-plus requests a day, and the current Pro model has not been free since April 2026. Two catches matter: free-tier data can be used to train Google's models, and enabling billing removes the free allowance rather than adding to it. ### Which is cheaper, the Gemini API or the ChatGPT API? For high-volume work, Gemini's Flash models usually undercut OpenAI's cheap tier: Gemini 2.5 Flash is $0.30 input against GPT-5.4-mini at $0.75. But reasoning tokens and context length can flip a base-rate comparison, so benchmark both on your real prompts. See our [ChatGPT API pricing](https://geotoolbox.ai/blog/chatgpt-pricing) guide for the full numbers. ### Why is my Gemini API bill higher than the token price suggests? Almost always thinking tokens, which bill at the output rate even though you never see them, plus grounding charges and the 200K context cliff on Pro models. None of the three shows up in a simple per-token estimate. ### What is the 200K context cliff? On the Pro models, once a single prompt crosses 200,000 tokens, the input rate roughly doubles and output rises with it. Gemini 3.1 Pro input goes from $2.00 to $4.00 per million. Retrieval pipelines cross it easily. ### Developer API or Vertex AI, which is cheaper? Both charge the same per-token rates. Vertex wins at high sustained volume through Provisioned Throughput discounts, but it adds Google Cloud costs like egress and idle endpoints. For most projects the Developer API is cheaper and simpler. ### Does the free tier train on my data? Yes. On the free tier, Google may use your inputs and outputs to improve its models, and human reviewers may read them. Paid usage on the Developer API and Vertex AI is excluded from training. ## Sources - Gemini Developer API pricing - Google (per-model token rates, service tiers, media meters) - `ai.google.dev/gemini-api/docs/pricing` - Gemini API billing and usage tiers - Google (tier caps, spend controls) - `ai.google.dev/gemini-api/docs/billing` - Gemini API additional terms - Google (free-tier data-use policy) - `ai.google.dev/gemini-api/terms` - Gemini API model deprecations - Google (shutdown and migration dates) - `ai.google.dev/gemini-api/docs/deprecations` - Vertex AI generative AI pricing - Google Cloud (rate parity, caching discounts) - `cloud.google.com/gemini-enterprise-agent-platform/generative-ai/pricing` - OpenAI API pricing - OpenAI (comparison rates) - `developers.openai.com/api/docs/pricing` --- ## Gemini Omni: What It Is, Pricing, and How to Get Access (2026) > Gemini Omni is Google's any-input-to-video AI model. What Omni Flash does, which Google AI plans include it, API pricing at $0.10/sec, and what breaks. - Canonical: https://geotoolbox.ai/blog/gemini-omni - Published: 2026-07-02 · Updated: 2026-07-19 Gemini Omni went from I/O stage demo to a model you can bill against in six weeks, and most of what ranks about it is already out of date. Here's what Gemini Omni is, which Google AI plans include it, what the API costs, and where the sharp edges are, verified against Google's launch and pricing documentation, and against what users have hit in practice since, as of July 19, 2026. ## What Is Gemini Omni? **Gemini Omni** is Google DeepMind's "any-to-any" model family: you feed it any mix of text, images, audio, and video, and it generates a finished video. Google announced the family at [Google I/O 2026](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni/) with the tagline "create anything from any input, starting with video," and began rolling the first member, **Gemini Omni Flash**, out to consumers the same day, May 19, 2026. Video is only the starting point. Google says image and audio output are on the Omni roadmap, which is why the family carries the "anything from anything" framing rather than a video-generator label.
![How Gemini Omni Flash works: mixed inputs, one model, a 10-second video, and the six official ways in — Gemini app, Google Flow, Google Vids, YouTube Shorts, AI Studio, and the Gemini API.](/blog/gemini-omni/gemini-omni-inputs-outputs-access-diagram.png)
One model takes any mix of inputs and returns a finished clip. Six official ways to reach it.
The internal pitch is simple: Nano Banana, but for video. Where Nano Banana handles image generation and editing inside the Gemini app, Omni does the same for moving pictures, drawing on what Google describes as Gemini's world knowledge of physics, history, and science. Google positions it as a native multimodal AI model rather than a text-to-video engine with adapters bolted on. And despite the name, this is not a GPT-4o-style live voice mode: "Omni" describes what the model accepts, not how you talk to it. One point worth clearing up, because early coverage muddied it: Gemini Omni is an official Google product line, with its own [DeepMind model page](https://deepmind.google/models/gemini-omni/) and API. Several explainers written before I/O treated "Gemini Omni" as a community nickname for Gemini-plus-Veo pipelines. That framing is now wrong, and some of those pages still rank. Omni sits inside the wider Gemini stack alongside the chat models, Veo, Imagen, and Lyria. If you want the full map of what Google ships under the Gemini name, we broke down [the wider Gemini ecosystem](https://geotoolbox.ai/blog/what-is-gemini) separately. ## Gemini Omni vs Veo: What Happened to Veo? The short version: **Omni replaces Veo inside the Gemini app**, and Veo carries on as the high-fidelity specialist everywhere else. Google's own product FAQ states it plainly: Gemini Omni is the newest video generation and editing model, and it takes over from Veo 3.1 as the default when you ask the Gemini app for video. As the rollout completes, video prompts in the app route to Omni Flash instead. Veo is not dead, though. The DeepMind lineup still lists Veo as a separate specialized model, and the Gemini API sells both. The practical split looks like this: Veo 3.1 is the realism engine, generating up to 4K broadcast-quality clips. Omni Flash tops out at 720p but accepts any input mix and lets you edit the result by talking to it. Here's the way to think about it. Veo is the cinematographer you hand a finished shot list. Omni is the editor sitting next to you who keeps the whole scene in its head while you change your mind, turn after turn. The cost gap runs the same direction: standard Veo 3.1 output costs four times as much per second as Omni Flash. ## What Gemini Omni Flash Can Do The headline capability is **conversational video editing**. Every instruction builds on the last one, and the scene keeps its continuity. DeepMind's demo sequence shows the pattern: "Transport the violinist to the image environment," then "Make the violin invisible," then "Change the camera angle to be over the violinist's shoulder," three plain-language turns on one shot. In those demos, characters hold their faces and clothing across edits without re-uploading references each turn. The reference system is the second pillar. In the consumer apps you can feed Omni Flash up to five reference photos, one video clip, and a text brief in one prompt, and it merges them into a single coherent output. Audio support starts narrow: voice references only, including **Avatars**, a digital version of you that looks and sounds like you in generated clips. Google gates avatar creation behind an onboarding flow intended to stop people from cloning someone else, and avatars are bound to the account holder's own likeness. From July 16, 2026, Google extended personal avatars into Google Vids: upload a selfie and a short voice recording, type a script, and your avatar delivers it. That feature is English-only, restricted to users 18 and over, and unavailable in the European Economic Area, Switzerland, and the UK. **Native audio** ships with every clip: sound is generated in the same pass as the video, so you get synchronized dialogue and effects rather than silent footage you dub later. Multi-turn **character consistency** rounds out the pitch. On quality, the benchmark numbers are strong, and for once Google published the methodology. On [DeepMind's evaluations](https://deepmind.google/models/gemini-omni/), human raters preferred Omni's edits over rival models across 504 side-by-side editing examples, and ranked it first for overall preference and instruction following on MovieGenBench's 1,003 text-to-video prompts. On the VBench image-to-video test (355 pairs), Omni Flash tied with Grok-Imagine-Video and Kling. Third-party signal agrees: [VentureBeat reported](https://venturebeat.com/technology/googles-gemini-omni-flash-hits-the-api-turning-enterprise-video-production-into-a-conversation) Omni Flash at number one on LMArena's Text-to-Video Arena with a score of 1527. Now the part the launch posts skip. Power users are split. In the 323-point [Hacker News thread](https://news.ycombinator.com/item?id=48196609) on the announcement, commenters picked apart the marketing demos: in the marble-physics showcase, "the marble jumps up for no reason." One heavy Seedance user was blunter: > "I've probably spent a couple grand on Seedance 2 to date, and I can't find anything google omni flash does better than Seedance from running a handful of samples through the system." The other recurring complaint is the content filter, and it has grown into the dominant theme of Google's own developer forum since launch. Our read of the tester consensus: Omni's real edge is editing and manipulating footage conversationally, not raw text-to-video quality, and "follows real-world physics" is a goal, not a guarantee. ### The Filter Problem Is Really a Billing Problem Weeks of actual use have surfaced something launch coverage did not: **a generation blocked by the safety filter still costs you credits.** Google's May fix exempted technical failures from app quota, but a policy block is a different path, and it does not trip the refund logic, so the failure is billed exactly like a success. One user's prompt for a porcelain statue in an orbital museum was flagged as potentially harmful and, in their words, "the 30 AI credits were not refunded." Another described a request for a green cinematic color grade being rejected instantly. The thread reporting this opened on May 20, 2026 and was still collecting replies in mid-July with no response from Google. The framing matters. This reads as a content-moderation argument and it is not one. Reasonable people disagree about where a filter should sit. Almost nobody thinks you should pay full price for an output you were never given. A related single report, worth less weight but the same root cause, describes a video-to-video edit returning the completely unmodified input while charging full cost, because a video came back and so nothing registered as a failure. ### The "Prominent People" Filter and the Flow-vs-App Split The best-documented bug in the window is a safety filter misfiring badly. Since late June, paying subscribers have reported that Google Flow's "prominent people" filter blocks **original fictional characters**, and in at least one case a user's own avatar after three prior successes. The error reads: "This prompt might violate our policies about generating prominent people." It survives prompt rewrites and character renaming. One user reported production halted for over a week; another, more than 20 days, and said they downgraded their plan and moved to a different engine. The workaround, which no launch coverage mentions, is a surface split: **the same prompts frequently succeed in the Gemini app while failing in Google Flow.** If you are blocked in Flow, try the app before you rewrite the prompt. That sits awkwardly beside the product direction. Google's July 16 update leans further into personal avatars while some paying users report being blocked from generating their own. ### Watermarks: Two of Them, and Only One Comes Off Google documents an invisible SynthID watermark on every generated clip. Users additionally report a **visible Gemini logo on the frame**, which is the one that actually blocks commercial work — as one put it, it "prevents it from being used in anything, from an Instagram post to a feature film." Google has not documented the visible mark, so treat it as a widely reported user observation rather than a published spec, but a small market of removal tools targeting Omni specifically has already appeared, which tells you how real it is. Worth knowing if you are tempted: those tools strip the visible overlay via reverse alpha blending. **SynthID survives, because it lives in the pixel values rather than as an overlay.** Removing the logo does not make a clip untraceable. ### Conversational Editing Holds for About Three to Five Turns The multi-turn editing is the genuine differentiator, and it has a shorter runway than the marketing implies. Users consistently report characters and environments being redesigned between turns, objects appearing mid-scene, and face consistency breaking down. One independent test put the reliable ceiling at four turns with drift beginning at the fifth. Google's own guidance of roughly three sequential edits is the honest number to plan against. That is still useful, and it is still ahead of the alternatives; it is just not the unbounded conversation the demos suggest. ## Specs and Current Limits The spec sheet, as the model ships today:
SpecGemini Omni Flash at launch
Model IDgemini-omni-flash-preview (public preview)
OutputVideo with native synchronized audio
Resolution720p, in 16:9 or 9:16
Clip length3 to 10 seconds at 720p. Longer durations were announced as coming; still unavailable as of July 19, 2026
InputsAny mix of text, images (up to 5 reference photos), video, and voice references
ProvenanceSynthID watermark on every clip; C2PA Content Credentials on Gemini app, Flow, and YouTube output, verifiable in the Gemini app
Age gate18+
The API carries a longer list of caveats. Per [Google's developer launch post](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/) and the [API docs](https://ai.google.dev/gemini-api/docs/omni): audio reference uploads and scene extension are still unsupported as of July 19, 2026 (the live docs state "uploading audio references is unsupported" and "video extension and video interpolation are not supported"), video references up to 3 seconds are "accepted by the API schema but are not correctly processed," multi-video reasoning and video interpolation are unsupported, and character consistency degrades on scene changes and panning shots. English is the only evaluated language so far. App-side, the limits are fuzzier by design. Gemini subscriptions use [compute-based usage limits](https://support.google.com/gemini/answer/16275805) that refresh every 5 hours under a weekly cap, with AI Plus at 2x standard limits, Pro at 4x, and Ultra at 5x or 20x Pro depending on the plan. Google publishes no per-video quota, and video generation burns compute fast. After the May 17 limits overhaul, subscribers filled Google's forums with complaints that one or two Omni generations emptied a full 5-hour window. Google's Gemini app VP Josh Woodward [acknowledged the bug by late May](https://www.androidauthority.com/gemini-usage-limit-changes-3672488/): failed requests no longer count against Gemini app quotas, and Ultra subscribers had their Omni video allowance doubled. Note the scope, because it is narrower than it sounds. That fix covers technical failures against the app's usage window. It does not cover generations stopped by the safety filter, and it does not cover Google Flow credits, which is where the complaints below are still coming from. ## How to Get Gemini Omni There are six official doors in, and one of them is free.
Access pathWho gets itCost
Gemini appGoogle AI Plus, Pro, and Ultra subscribers, globally, 18+From $4.99/mo (AI Plus)
Google FlowSame subscribers; Flow AI credits: 200/mo on Plus, 1,000 on Pro, 10,000 to 25,000 on UltraIncluded in plan
YouTube Shorts + YouTube CreateYouTube users as the rollout reaches themFree
Google AI Studio + Gemini APIDevelopers, public preview$0.10 per second of video, paid tier only
Google VidsWorkspace Business, Enterprise, Education Plus and Nonprofits tiers, plus consumer AI Pro and Ultra. Rolling out from July 16, 2026Included in plan
Gemini Enterprise Agent PlatformEnterprise customersEnterprise terms
The cheapest paid route is **Google AI Plus** at $4.99 a month (cut from $7.99 in June 2026), which includes Omni Flash with the lowest usage ceiling. Pro at $19.99 raises the limits and [the Flow credit pool](https://one.google.com/about/google-ai-plans/). We keep the full tier-by-tier breakdown current in our [Gemini pricing breakdown](https://geotoolbox.ai/blog/gemini-pricing), including what Ultra actually costs. Two access restrictions catch people out. First, per the official API documentation, editing uploaded videos is not available in the European Economic Area, Switzerland, or the UK, and Google's app help page adds some US states to that list; images containing minors are blocked from upload in the EEA, Switzerland, and the UK. Second, business access has its own rules: personal accounts need a Google AI plan, while [work and school accounts need a qualifying Workspace license](https://support.google.com/gemini/answer/16126339), a distinction that filled Google's support forum with locked-out business users in the launch window. There is no official Gemini Omni APK, login portal, or desktop app. "Gemini omni apk" and "gemini omni app download" searches lead to squatter sites, and as of July 2026 unofficial domains still rank on page one for the model's name. Consumer access runs through Google's surfaces listed above; developers can also reach the model through licensed API platforms, never through download portals. ## Gemini Omni API Pricing Developer access arrived on June 30, 2026, when Google [brought Omni Flash to the Gemini API and AI Studio](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite/) in public preview, priced at **$0.10 per second of generated video**. A 10-second clip, the current maximum, costs about a dollar. Under the hood that dollar is token math: video output is billed at [5,792 tokens per second of 720p footage](https://ai.google.dev/gemini-api/docs/pricing) against a $17.50 per million video-output-token rate, with input at $1.50 per million tokens across text, image, video, and audio. Here's how that sits against the rest of Google's video lineup at 720p:
Model720p, per secondPositioning
Veo 3.1 Lite$0.05Cheapest, 1080p max
Gemini Omni Flash$0.10Any-input generation + conversational editing
Veo 3.1 Fast$0.10Speed-tier realism, up to 4K at $0.30
Veo 3.1$0.40Full realism tier, up to 4K at $0.60
The multi-turn editing runs on Google's **Interactions API**, a stateful interface that carries the previous video and its references forward so you can stack up to three sequential edits with the session context carried forward. It is the same multimodal session layer Nano Banana uses, which is what makes the two models chainable. Google's intended pattern is chaining: generate stills with **Nano Banana 2 Lite**, the sibling image model launched alongside it at $0.034 per 1K-resolution image with 4-second latency, then hand them to Omni Flash to animate and refine. Google shipped three remixable demo apps (Anywhere, Space Lift, and Omni Product Studio) to show the image-to-video pipeline end to end. For enterprise buyers, Omni Flash is live in the Gemini Enterprise Agent Platform, and this API release is the moment the model stopped being a consumer toy. As VentureBeat put it, the missing programmatic interface was "the catch" at I/O; the API rollout is what puts conversational editing in front of the marketing and training teams that produce most corporate video. One naming trap if you build on Google Cloud: searching for "Vertex AI Omni" returns false negatives, because Google renamed Vertex AI to the Gemini Enterprise Agent Platform in April 2026. It is the same product. `gemini-omni-flash-preview` is documented there, at preview stage. ## How Gemini Omni Stacks Up Against Rivals The competitive backdrop shifted twice in spring 2026. OpenAI [shut down the Sora app on April 26, 2026](https://the-decoder.com/openai-sets-two-stage-sora-shutdown-with-app-closing-april-2026-and-api-following-in-september/), with the API following in September, taking the most famous consumer video generator off the board. Three weeks later, Google announced Omni. That leaves a different rival at the top of power users' rankings: **ByteDance's Seedance 2.0**, which testers consistently cite for raw generation quality, higher resolution output, and bigger reference budgets per generation. The Hacker News verdict quoted earlier came from someone who had spent thousands of dollars on Seedance and saw no reason to switch. The rest of the field: **Kling** and **Grok-Imagine-Video** tied Omni Flash on DeepMind's own image-to-video benchmark, so treat "leading results" claims from any of the three with that context. We covered xAI's entry separately in our [Grok Imagine](https://geotoolbox.ai/blog/grok-imagine) review. Runway remains the professional editing suite of the group. Omni's genuine differentiators are narrower than the marketing but real: conversational editing with scene memory, the stateful Interactions API workflow, and a $0.10-per-second price that undercuts most premium rivals. Native audio in a single pass helps, though it is no longer unique, and former Sora users rate Omni's generated voices as robotic. Where it loses today: resolution (720p vs 1080p-4K elsewhere), clip length, stylized output, and, by heavy-user consensus, raw text-to-video fidelity. ## What Gemini Omni Means for Brands and GEO Video is becoming an answer surface, and Omni accelerates that. When we ran the LLM citation data for this exact topic through DataForSEO's mentions index (July 2026, tested), YouTube was the single most-cited domain in Google's AI answers about Gemini Omni: 483 of 737 tracked citations, ahead of Google's own blog. AI engines already lean on video pages to answer questions; a model that lets anyone produce credible product video at $1 per clip will flood that surface. Three practical takeaways for anyone managing a brand's AI visibility: **Provenance becomes a trust signal.** Every Omni clip carries SynthID, consumer-surface output adds C2PA Content Credentials, and Google is wiring verification into Search, Chrome, and the Gemini app. Brands publishing real footage should expect provenance signals to start separating them from synthetic filler. **Watch video citations, not just text.** In our experience at GEO Toolbox, teams tracking AI visibility monitor text answers and skip video entirely, even though YouTube already dominates citations on queries like this one. If your competitors' clips get cited in AI answers and yours don't exist, that gap won't show up in any keyword-ranking report. **The ecosystem is a distribution channel.** [Picsart is putting Omni Flash in front of 130 million creators](https://deepmind.google/models/gemini-omni/), with Artlist, OpusClip, and Higgsfield running similar integrations. Branded video volume is about to spike, and the same brand-kit consistency that Omni's partners sell is what keeps machine-generated brand mentions on-message. The family is also just getting started: Google has teased image and audio output, launch coverage points to a heavier Omni model above Flash, and the drip-release cycle looks like the one we tracked with [Gemini 3.5 Pro](https://geotoolbox.ai/blog/gemini-3-5-pro). If Omni content starts answering questions in your category, the playbook for [getting cited in Gemini](https://geotoolbox.ai/blog/gemini-seo) applies to your video the same way it applies to your pages. ## Where This Goes Next Omni Flash has only been a developer product since June 30, and Google is shipping against a public roadmap: longer clips, audio reference support, and image and audio output, with a heavier Omni model reported on the way. Worth noting how little of that roadmap has landed: as of July 19, 2026, none of it has. `gemini-omni-flash-preview` is still the only Omni model ID, still tagged preview rather than GA, still video-output-only, still capped at 10 seconds. Two things did ship, and neither is on the roadmap list: developer logs on the Interactions API on July 6, aimed squarely at the opaque-rejection complaints above, and Omni landing in Google Vids with personal avatars on July 16. The model has not moved; its distribution has. Expect the spec table to age in weeks, not years. We update this page as the family grows. Meanwhile, the searches AI engines answer about your brand are already being fed by whoever publishes first, in text and now in video. If you want to know where you stand before that wave hits, you can [check how AI systems see your site](https://geotoolbox.ai/tools/ai-readiness) with GEO Toolbox's free AI readiness scan; it takes about a minute and shows what the crawlers behind these answers can actually reach. ## FAQ ### Is Gemini Omni released? Yes. Google announced the Omni family at I/O 2026 on May 19 and began rolling Gemini Omni Flash out to Google AI subscribers in the Gemini app the same day, then opened developer access via the Gemini API on June 30, 2026. It remains labeled a preview on the API side. ### Is Gemini Omni free? Partly. Omni Flash video generation is rolling out at no cost inside YouTube Shorts and the YouTube Create app. Using it in the Gemini app or Google Flow requires a Google AI Plus ($4.99/mo), Pro, or Ultra subscription, and API use is paid-tier only. ### How much does the Gemini Omni API cost? $0.10 per second of generated 720p video, billed as 5,792 video tokens per second at $17.50 per million output tokens. A maximum-length 10-second clip costs about $1, the same per-second rate as Veo 3.1 Fast. ### What is the difference between Gemini Omni and Veo? Omni replaces Veo as the video model inside the Gemini app and focuses on any-input generation and conversational editing at 720p. Veo 3.1 continues separately as Google's realism specialist, generating up to 4K clips at up to $0.60 per second via the API. ### Can you use Gemini Omni with a Google Workspace account? Yes, with the right license. Google's video generation help page says work and school accounts need a qualifying Workspace license, while personal accounts need a Google AI plan (Plus, Pro, or Ultra). Many Workspace users reported being locked out in the launch window before licensing caught up, and developers on any account type can use the paid Gemini API. ### Why does Gemini Omni say my prompt violates the prominent people policy? This is a known filter misfire that has been affecting paying Google Flow users since late June 2026, and it blocks original fictional characters and even users' own avatars, not just real public figures. Renaming the character or rewriting the prompt generally does not clear it. The workaround users report is a surface split: the same prompt often succeeds in the Gemini app while failing in Flow, so try the app before you rewrite anything. ### Do I get my credits back if Omni blocks my video? No. A generation stopped by the safety filter is billed the same as a successful one, because a policy block does not trigger the refund path. Users have been reporting this on Google's own developer forum since May 2026 without resolution. Budget for it: if you are running prompts that sit anywhere near the filter's boundaries, some share of your credits will buy you nothing. ### How do I remove the Gemini Omni watermark? There are two watermarks, and only one is removable. Users report a visible Gemini logo on the frame, and third-party tools do strip it. The invisible SynthID watermark that Google embeds in every clip lives in the pixel values rather than as an overlay, so it survives removal. Taking the logo off does not make a clip untraceable or unattributable. ### How many edits can Gemini Omni handle before it breaks? Plan for about three sequential edits, which matches Google's own guidance. Users and independent testers report drift starting around the fourth or fifth turn: characters get redesigned, objects appear mid-scene, and faces stop matching. Conversational editing is genuinely Omni's best feature, but it has a shorter runway than the demos suggest. ### Can you make Gemini Omni videos longer than 10 seconds? Not as of July 19, 2026. Clips run 3 to 10 seconds at 720p and 24fps. Google announced longer durations as coming and described the 10-second cap as a rollout decision rather than a model limit, but nothing longer has shipped. ### Is Gemini Omni available on Vertex AI? Yes, though the name is the problem. Google renamed Vertex AI to the Gemini Enterprise Agent Platform in April 2026, so searching for "Vertex AI Omni" turns up nothing useful. The model is documented there as `gemini-omni-flash-preview`, at preview stage. ### Can you use Gemini Omni videos commercially? Generally yes. Google's terms do not claim ownership of generated output, and commercial use is allowed within its content policies. Every clip carries the invisible SynthID watermark no matter where it was made (users report a visible badge on consumer-app clips too), and purely AI-generated footage may not qualify for copyright protection on its own. ## Sources - Introducing Gemini Omni - blog.google - `blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni` - Start building with Nano Banana 2 Lite and Gemini Omni Flash - blog.google - `blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-omni-flash-nano-banana-2-lite` - Gemini Omni Flash API documentation - ai.google.dev - `ai.google.dev/gemini-api/docs/omni` - Gemini Omni model page and benchmarks - Google DeepMind - `deepmind.google/models/gemini-omni` - Gemini Apps limits and upgrades for Google AI subscribers - Google Support - `support.google.com/gemini/answer/16275805` - Generate videos with Gemini Apps - Google Support - `support.google.com/gemini/answer/16126339` - Gemini Developer API pricing - ai.google.dev - `ai.google.dev/gemini-api/docs/pricing` - Google AI plans - Google One - `one.google.com/about/google-ai-plans` - Google's Gemini Omni Flash hits the API - VentureBeat - `venturebeat.com/technology/googles-gemini-omni-flash-hits-the-api-turning-enterprise-video-production-into-a-conversation` - Gemini Omni discussion - Hacker News - `news.ycombinator.com/item?id=48196609` - Google may have fixed the issue exhausting Gemini usage limits - Android Authority - `androidauthority.com/gemini-usage-limit-changes-3672488` - OpenAI sets two-stage Sora shutdown - The Decoder - `the-decoder.com/openai-sets-two-stage-sora-shutdown-with-app-closing-april-2026-and-api-following-in-september` --- ## Google AI Mode SEO: How to Get Cited and Track It > How Google AI Mode picks and cites pages, what actually changes for SEO, how to measure visibility Search Console barely breaks out, and what you can control. - Canonical: https://geotoolbox.ai/blog/google-ai-mode-seo - Published: 2026-07-02 · Updated: 2026-08-05 Google AI Mode SEO is a different job from ranking. In the conversational tab, you are not competing for a spot in a list of links, you are trying to be one of the few sources the generated answer cites, and you often cannot see whether you made it. This guide covers how AI Mode picks pages, what actually changes in your SEO, how to measure visibility Search Console barely breaks out, and which levers you genuinely control. ## What Google AI Mode Is, and How It Differs from AI Overviews Google AI Mode is the conversational, Gemini-powered search experience you open as its own tab. Instead of a page of ten blue links, you get a generated answer you can follow up on, with a handful of cited sources attached. It [surpassed one billion monthly users](https://blog.google/products-and-platforms/products/search/search-io-2026/) a year after launch, per Google at I/O 2026, and it is now free and broadly available rather than the US-only, opt-in Labs experiment it started as. Most of the confusion here comes from mixing up three names. Search Generative Experience (SGE) was the Labs prototype. Google retired that name and expanded the generative summary experience into what shipped broadly as AI Overviews. AI Mode is the [separate, full conversational surface](https://geotoolbox.ai/blog/what-is-google-ai-mode) that arrived after. They are related, they share Google's index, but they are not the same product and they do not behave the same way. The distinction matters for SEO because the two surfaces select and show sources differently.
 AI OverviewsAI Mode
Where it livesA summary box on the normal results page, above the blue linksA separate tab you open; it replaces the results page
How it triggersAuto-injected when Google judges it additive to classic SearchUser-initiated; you choose to search in it
InteractionMostly read-and-scroll; less conversational than AI ModeMulti-turn conversation with follow-ups
Query fan-outLighter; a few subtopicsHeavier; many concurrent sub-searches per question
SEO goalBe one of the cited sources in the boxBe a cited source inside the synthesized answer
If your focus is the summary box on the standard results page, that is a related but separate playbook, covered in [how to get cited in Google AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo). This piece is about the conversational tab, where the fan-out runs deeper and measurement gets harder. For a plain-English definition you can point a colleague to, the [Google AI Mode glossary entry](https://geotoolbox.ai/glossary/google-ai-mode) covers the basics. ## How AI Mode Picks the Pages It Cites AI Mode does not run your query and rank pages. It takes your question, breaks it into subtopics, and runs many searches at once. Google calls this a [query fan-out technique](https://blog.google/products/search/ai-mode-search/), "issuing multiple related searches concurrently across subtopics and multiple data sources," then it synthesizes one answer and attaches the sources that supported it. It now runs this as an agentic loop rather than a single pass: AI Mode builds a search plan, fires off the sub-queries, evaluates what comes back, and adjusts before it commits to an answer. It is the same fan-out architecture as AI Overviews, with deeper reasoning layered on top, which is why a page has to earn its place against a moving set of sub-questions rather than one fixed query. Picture a search like "best CRM for a 20-person B2B team with HubSpot integration under $500 a month." AI Mode may quietly fan that into separate searches for CRM options at that team size, HubSpot integrations, pricing tiers, and security requirements. A different page can be pulled in for each strand. Your page gets cited because it answered one of those hidden sub-questions well, not because it held position one for the sentence the user typed. If you want the mechanics in depth, we covered [query fan-out](https://geotoolbox.ai/blog/query-fan-out) separately. ![Google AI Mode turns one question into many sub-searches through query fan-out, then writes one answer citing the pages whose passages best fit each sub-question.](/blog/google-ai-mode-seo/query-fan-out-selection.png) Two consequences follow, and both shape the rest of this guide. First, treat the passage as the unit of selection. Google documents supporting links and query fan-out rather than a formal passage-ranking rule, but in practice AI Mode lifts the specific chunk that answers a strand. A clear, self-contained passage under a descriptive heading behaves like the thing that gets cited, not the page as a whole. Second, ranking and citation have come apart. You still need to be indexed and eligible for a snippet, but you do not need to hold position one. Google says AI features surface a "wider and more diverse set of helpful links" than classic search, so a page sitting on page two for the head term can still be cited when its passage best answers a fanned-out strand. That is the most encouraging point for a smaller site. The citation set is also unstable: a Semrush study, reported by [iPullRank](https://ipullrank.com/everything-we-know-about-ai-overviews), found 91% of URLs cited in AI Overviews were dropped at some point, so a page cited today can be gone next week (that figure is from AI Overviews, the closest measured proxy). And Google cites itself heavily, with SE Ranking's analysis of 1.3 million AI Mode citations putting [Google.com as the most-cited domain at 17.42%](https://seranking.com/blog/google-links-in-ai-mode-answers/), more than the next several combined. The winnable ground is the specific sub-question, not the head term Google already owns. To understand why a page qualifies at all, it helps to know [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) under the hood. ## Does AI Mode Need a Different SEO Strategy? Google's own answer is blunt. Its documentation states there are ["no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary,"](https://developers.google.com/search/docs/appearance/ai-features) and that to be eligible a page "must be indexed and eligible to be shown in Google Search with a snippet." AI Mode is not a separate index you submit to. It reads from normal search. So the honest read is that this is not a new discipline bolted onto SEO. It is SEO, with the prerequisite made non-negotiable and a few emphases shifted. If Google cannot crawl, render, index, and trust your page, it will not become a reliable AI Mode source. There is no shortcut that skips being indexed and genuinely relevant. Anyone selling you a way to "optimize for the AI" that bypasses being indexed and relevant is selling a rebrand. That is the practical answer to the [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) debate: the label is new, the foundation is not. What genuinely shifts is narrower, and worth naming so you spend effort in the right place: - Visibility becomes binary. You are cited or you are not, rather than sitting at position seven where a determined user might still find you. - Sub-intent coverage matters more than the head keyword, because the fan-out rewards pages that answer the strands around a topic. - Extractability matters more than it used to, because a passage has to stand on its own to be lifted into an answer. - Corroboration matters more, because AI Mode leans on sources that agree with each other across the web. Everything else you already do well, keep doing. The next section turns those four shifts into concrete moves. ## How to Optimize for Google AI Mode Here is the checklist, highest-impact first. 1. **Cover the fan-out on one deep page.** List the sub-questions a real user would ask around your topic, then answer each one on the same page under its own heading. If a page answers the pricing strand, the integration strand, and the "who is this for" strand, it is eligible for more of the fan-out searches than a thin page that only defines the term. This is coverage on one page, not a doorway page per sub-question. 2. **Write answer-first, self-contained passages.** Lead each section with the direct answer in a sentence or two, then support it. AI Mode lifts passages, so a chunk that only makes sense after reading the whole article is hard to cite. Descriptive headings that mirror the question help the model match your passage to a strand. 3. **Use structured data honestly.** Add schema that matches your visible content: Article, FAQPage, Product where it applies. It helps Google parse and trust what a section is, and it is required for some rich results. It is not a citation switch, and Google does not treat it as a ranking factor. We go deeper on where it helps and where it is oversold in [schema markup for AI](https://geotoolbox.ai/blog/schema-markup-for-ai). While you are checking files, skip the effort on speculative ones: [Google does not use or document llms.txt for AI Mode](https://geotoolbox.ai/blog/llms-txt), which pulls from the normal search index. 4. **Build corroboration and entity signals off-site.** AI Mode favors claims that hold up across sources. Consistent brand and author identity, credible mentions on the places these answers already pull from, and a clean [entity footprint](https://geotoolbox.ai/blog/entity-seo) all raise the odds you are chosen when a strand needs a trusted source. The [broader AEO checklist](https://geotoolbox.ai/blog/aeo-best-practices) covers the on-page and off-page basics that still underpin all of this. 5. **Show real experience.** Firsthand tests, original data, named authors, and evidence a competitor cannot copy make a passage harder to replace with a generic one. This is E-E-A-T doing the same job it always did, now as a tiebreaker for citation. 6. **For products and local, feed the right graph.** AI Mode does not only read the web index. Its fan-out also pulls from the Knowledge Graph, the Shopping Graph, and live sources, so for e-commerce and local queries the real levers are an accurate Merchant Center feed and an up-to-date Google Business Profile, not just blog passages. In our experience running scans across the geotoolbox engine set, the pages that get pulled into AI Mode answers are rarely the flashiest. They are the ones that answer a narrow question cleanly, in a passage you could lift out and paste into a reply without editing. Write for that. ## The Traffic Reality: Impressions Up, Clicks Down Be honest with yourself and your stakeholders about clicks. The trend is real and it predates AI Mode. Around [58.5% of US Google searches already ended without a click](https://sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches-only-374-clicks-go-to-the-open-web-in-the-eu-its-360/) in 2024, before AI Mode was widespread. AI features push that further. Ahrefs, studying 300,000 keywords, found the presence of an AI Overview correlated with a [34.5% lower click-through rate](https://ahrefs.com/blog/ai-overviews-reduce-clicks/) for the top-ranking page. BrightEdge's one-year data showed [impressions up more than 49% while click-throughs fell about 30%](https://www.brightedge.com/news/press-releases/one-year-google-ai-overviews-brightedge-data-reveals-google-search-usage). Inside AI Mode the gap looks sharper still: Seer Interactive, analyzing roughly 25 million impressions, found that 93% of AI Mode queries produced no outbound click, versus about 43% for AI Overviews. One caveat worth stating plainly: AI-Mode-specific click data is still thin, and most of the numbers above come from AI Overviews studies, so treat the AI Mode figures as early signal rather than settled fact. The direction, though, is consistent. That is the "great decoupling": you are shown more and clicked less. Some of the industry reaction is stark. Speaking to [Technology Magazine](https://technologymagazine.com/articles/how-googles-new-ai-mode-could-devastate-web-traffic-seo), Lily Ray of Amsive warned that making AI Mode the default "is going to have a devastating impact on the internet," and Barry Adams of Polemic Digital expects click volume to the web to roughly halve. Google's framing is that the clicks you do get are [worth more](https://developers.google.com/search/docs/appearance/ai-features), with users spending more time on site, and its own executives argue the web is thriving. You do not have to pick a camp to act on this. The move is to stop grading yourself only on click volume for informational queries and start counting the value AI Mode actually creates: presence in the answer, brand recall, and higher-intent visits. NerdWallet is the case in point. Its [monthly unique users fell about 20% year over year while quarterly revenue rose 37%](https://investors.nerdwallet.com/news-releases/news-release-details/nerdwallet-reports-fourth-quarter-and-full-year-2024-results) in Q4 2024. That revenue came from other parts of the business, not from AI Mode, but it shows traffic and revenue can move in opposite directions. The point stands only if you can see whether you are in the answer at all, which is where most teams hit a wall. ## How to Measure and Track AI Mode Visibility Start by knowing what Search Console can and cannot tell you, because this is where most of the confusion lives. In the main Performance report, AI Mode and AI Overview activity is [folded into the "Web" search type](https://developers.google.com/search/docs/appearance/ai-features) with no filter to isolate it. In June 2026 Google added a separate [generative AI performance report](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports), but read the fine print: it reports impressions only, not clicks, and it lumps AI Overviews and AI Mode together into one "generative AI features in Search" bucket you cannot split apart. (AI Overviews appearing in Discover gets its own separate report, not folded into this one.) It also rolled out in phases. So you can increasingly see that you appeared in an AI surface, but not which one, and not what it drove. Three more things break the old measurement habits: - There is no fixed top ten to track. AI Mode synthesizes an answer, so traditional rank trackers have nothing to report a position against. - Answers are personalized and non-deterministic. Two people asking the same question can see different sources, so checking once and seeing yourself proves little. - Attribution has been shaky. For a stretch in 2025, AI Mode links [carried a noreferrer attribute](https://searchengineland.com/googles-ai-mode-traffic-untrackable-455883) that hid the referrer in analytics; Google said it was unexpected and fixed it, but the episode is a reminder that AI-surface traffic is easy to misread. So you measure presence, not position, and you measure it across many samples rather than one lucky check. Practically that means tracking whether you are cited for a set of buyer questions, how your share of those answers trends over time, and whether branded search lifts as AI Mode mentions you. Our own writeups on [measuring AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) and [AI Overview trackers](https://geotoolbox.ai/blog/ai-overview-tracker) walk through the free and paid options. This is the gap geotoolbox was built for. Our scan runs your prompts across eight engines, and Google AI Mode is one of them, so you get a straight read on whether AI Mode cites you or a competitor for the questions your buyers actually ask, tracked across repeated scans so you are reading a trend rather than a single non-deterministic snapshot. Even that is a probe, not any one user's personalized session. It will not give you a clean click number, because no tool honestly can yet. It will tell you whether you are in the answer. ## What You Can Actually Control A lot of anxiety here comes from misunderstanding the levers, so here is what each one actually does. Googlebot governs whether Google can crawl your site for Search, and since AI Mode reads from the search index, that same crawl access is what makes you eligible for AI features. Two of these are real limits: `data-nosnippet` hides specific elements, and `max-snippet` caps snippet length. Be careful with the blunt ones. `nosnippet` and `max-snippet:0` make the page snippet-ineligible, and since AI features require a page to be eligible for a snippet, they effectively remove you from AI Mode citation rather than just trimming it. `noindex` removes the page from Search entirely. The most common mistake is reaching for Google-Extended to get out of AI Mode. It does not do that. [Google-Extended controls training and grounding](https://developers.google.com/search/docs/appearance/ai-features) in Google's other AI products, not your appearance in AI Overviews or AI Mode, which run on the normal search index through Googlebot. Block Google-Extended and you change what future models train on without opting yourself out of AI Mode or AI Overviews. If you want to understand the wider set of AI crawlers and what each one governs, we keep a running guide to [AI crawlers and how to control them](https://geotoolbox.ai/blog/ai-crawlers). There is a genuine opt-out, but it is a blunt one and it isn't available everywhere yet. Google began testing its Search Console generative-AI control on June 3, 2026 with a subset of UK properties, and started respecting the setting on June 17, 2026 under a binding Competition and Markets Authority conduct requirement. Since July 2026 it has also been appearing on accounts outside the UK, though the rollout remains partial with no announced completion date. Google says the toggle is not used as a ranking signal, and it does not cover the Gemini app. Where it's live, it keeps a site out of AI Overviews and AI Mode, and the tradeoff is total: you forfeit visibility and impressions from those surfaces while staying in classic results. For the vast majority of sites that is the wrong trade, because the traffic is leaving the open web regardless and being absent from the answer does not bring the click back. The better play is to be the source the answer cites. ## Frequently Asked Questions ### What is the difference between Google AI Mode and AI Overviews? AI Overviews is the summary box that appears above the blue links on a normal results page. AI Mode is a separate conversational tab that replaces the results page with a generated, multi-turn answer. Both are powered by Gemini and share Google's index, but AI Mode fans a question out into more concurrent searches and is where being a cited source, not a ranked link, is the whole game. ### Do I need a different SEO strategy for AI Mode, or is normal SEO enough? It is mostly the same SEO with the prerequisite made non-negotiable. Google states there are no additional requirements to appear, and a page must simply be indexed and eligible for a snippet. What shifts is emphasis: cover the sub-questions around a topic, write passages that stand on their own, and earn corroboration off-site. There is no path that skips being indexed and relevant. ### How do I track whether my site appears in Google AI Mode? You measure presence across many samples rather than a single position. Search Console now shows generative-AI impressions in a separate report, but it lumps AI Mode in with AI Overviews and shows no clicks, while its main report folds AI activity into the "Web" search type. Use a tool that samples your buyer questions across engines to see whether AI Mode cites you, and watch branded search for lift. ### Why did my impressions go up but my clicks go down? Because AI features show your link more often while answering the question in place, so fewer people click through. BrightEdge measured impressions up roughly 49% and clicks down about 30% across a year of AI Overviews. The fix is not to chase the lost clicks but to value the presence and the higher-intent visits that remain. ### Does schema markup or llms.txt help me get cited in AI Mode? Schema that matches your visible content helps Google parse and trust a section and is needed for some rich results, but it is not a citation switch or a ranking factor. Google does not use llms.txt for AI Mode, which retrieves from the normal search index, so that effort is better spent on indexable, well-structured content. ### Why doesn't ranking #1 get me cited in the AI answer? Because AI Mode selects passages that answer the sub-questions it fanned out, not the single page that ranks first for the typed query. Citation sets are also volatile and self-referential, with one study finding Google.com alone accounts for 17.42% of AI Mode citations. A page can rank first and still be absent from the answer if a narrower passage elsewhere fits a strand better. ### Can I opt out of AI Mode, and does blocking Google-Extended stop it? There is a Search Console control that removes a site from AI Overviews and AI Mode while keeping it in regular Search, and Google says it is not used as a ranking signal. Google began testing it on June 3, 2026 with a subset of UK properties and started respecting the setting on June 17, 2026 under a binding CMA conduct requirement; since July 2026 it has been appearing outside the UK too, with the rollout still partial and no completion date announced. But it forfeits all visibility and impressions from those AI surfaces, which is the wrong trade for most sites. Blocking Google-Extended does not opt you out: it governs training in Google's other AI products, while AI Mode appearance runs through Googlebot and the normal index. The searcher-side question, hiding Google's AI answers from your own results, is covered in [turning off AI Overviews](https://geotoolbox.ai/blog/how-to-turn-off-ai-overviews). ## Track Your AI Mode Visibility The uncomfortable part of AI Mode is how little you can see. You cannot check a rank, clicks are hidden, and the answer changes per person. What you can do is stop guessing. Fix the fundamentals that make a page citable, answer the questions around your topic in chunks that make sense on their own, and then measure presence instead of position. If you want to know whether Google AI Mode is citing you or your competitor for the questions your buyers ask, [run a scan across the eight engines](https://geotoolbox.ai/features/geo-scan) and see where you actually stand. geotoolbox treats AI Mode as one of those engines, so you get a real read on the answer, not a rank you can no longer track. ## Sources - Google Search Central: AI features and your website - `developers.google.com/search/docs/appearance/ai-features` - Google: Expanding AI Overviews and introducing AI Mode - `blog.google/products/search/ai-mode-search` - Google: A new era for AI Search (I/O 2026) - `blog.google/products-and-platforms/products/search/search-io-2026` - Google Search Central: Generative AI performance reports - `developers.google.com/search/blog/2026/06/gen-ai-performance-reports` - SparkToro: 2024 Zero-Click Search Study - `sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches-only-374-clicks-go-to-the-open-web-in-the-eu-its-360` - Ahrefs: AI Overviews reduce clicks - `ahrefs.com/blog/ai-overviews-reduce-clicks` - BrightEdge: One year of Google AI Overviews - `brightedge.com/news/press-releases/one-year-google-ai-overviews-brightedge-data-reveals-google-search-usage` - SE Ranking: Google links in AI Mode answers - `seranking.com/blog/google-links-in-ai-mode-answers` - iPullRank: Everything we know about AI Overviews (reporting a Semrush citation-volatility study) - `ipullrank.com/everything-we-know-about-ai-overviews` - NerdWallet Q4 and full-year 2024 results - `investors.nerdwallet.com/news-releases/news-release-details/nerdwallet-reports-fourth-quarter-and-full-year-2024-results` - Search Engine Land: Google AI Mode traffic untrackable - `searchengineland.com/googles-ai-mode-traffic-untrackable-455883` --- ## Claude Sonnet 5: Pricing, Benchmarks & Real-World Review (2026) > Claude Sonnet 5 is out. Pricing, benchmarks vs Sonnet 4.6 and Opus 4.8, the real per-task cost at high effort, and what the first independent reviews reveal. - Canonical: https://geotoolbox.ai/blog/claude-sonnet-5 - Published: 2026-06-30 · Updated: 2026-08-14 Claude Sonnet 5 launched June 30, 2026. Months of leaked rumors under the wrong codename preceded it, and the real model lands close to what those rumors promised: most of Opus 4.8's agentic ability, at a meaningfully lower price, with a few real caveats the launch posts gloss over. ## What Is Claude Sonnet 5? Claude Sonnet 5 is Anthropic's mid-tier Claude model, released June 30, 2026. It sits between the fast, low-cost Haiku 4.5 and the higher-tier Opus and Fable 5 models, and Anthropic calls it "the most agentic Sonnet model yet." Note that the Opus tier has moved on since this article was written. Anthropic released [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5) on July 24, 2026, replacing Opus 4.8 at the same $5 and $25 per million tokens. The Opus 4.8 comparisons below still describe the model Sonnet 5 launched against. That framing is specific, not marketing filler. Sonnet 5 can build a multi-step plan, decide which tools it needs (a browser, a terminal, a file editor), execute that plan with minimal hand-holding, and check its own output before handing back a result. Anthropic's own write-up describes testers asking it to investigate a bug. Unprompted, it wrote a reproducing test, implemented the fix, then stashed the change to confirm the bug actually came back without it, all in one pass. Sonnet 5 closes most of the gap to Opus 4.8 while staying at Sonnet-tier pricing. For many agentic and coding tasks, [Claude](https://geotoolbox.ai/blog/what-is-claude-ai) users now have a meaningfully cheaper option that doesn't feel like a downgrade. That doesn't hold at every effort setting, though, and the real cost math gets its own section below.
![Bar chart of SWE-bench Pro scores: Sonnet 4.6 at 58.1%, Sonnet 5 at 63.2%, and Opus 4.8 at 69.2%](/blog/claude-sonnet-5/benchmark-comparison.png)
On SWE-bench Pro, Sonnet 5 clears Sonnet 4.6 by five points but still trails Opus 4.8 by six. It closes most of the gap, not all of it.
## No, You're Not Misremembering "Sonnet 5" from Months Ago If "Claude Sonnet 5" sounds familiar, you're thinking of "Fennec." In early February 2026, [a Google Vertex AI error log exposed a model identifier](https://dev.to/marc0dev/claude-sonnet-5-fennec-leak-what-the-vertex-ai-logs-actually-show-3ho5), `claude-sonnet-5@20260203`, alongside that internal codename. It looked like proof a launch was imminent, and a wave of speculative coverage ran with it for months, including fabricated benchmark tables and at least one [April Fool's Day satire post](https://dev.to/best_codes/anthropic-just-dropped-claude-sonnet-5-and-the-benchmarks-are-kind-of-insane-3ppc) with invented scores like 92.4% on SWE-bench Verified. None of it came from Anthropic. What actually happened: [that leaked checkpoint shipped as Claude Sonnet 4.6](https://www.nxcode.io/resources/news/claude-sonnet-5-fennec-leak-2026) on February 17, 2026, not as Sonnet 5. Sonnet 5 itself launched four months later, on June 30, 2026, with an official announcement, a system card, and the numbers below. If you've been holding onto a "Fennec" spec sheet, throw it out, it was describing a different model. ## Claude Sonnet 5 Pricing [Claude Sonnet 5 launched with introductory pricing](https://www.anthropic.com/news/claude-sonnet-5) that Anthropic made permanent in August 2026, cancelling the planned increase:
PeriodInput (per million tokens)Output (per million tokens)
Introductory (announced through Aug 31, 2026)$2.00$10.00
Now permanent (increase to $3/$15 cancelled)$2.00$10.00
For comparison, Claude Opus 4.8 runs $5 per million input tokens and $25 per million output tokens (see our full [Claude pricing](https://geotoolbox.ai/blog/claude-pricing) guide). Sonnet 5's price is roughly 60% cheaper than Opus 4.8 on both ends. The tokenizer changed too, so the price-per-token comparison overstates the savings. Sonnet 5 runs on an updated tokenizer, the same change Anthropic introduced with Opus 4.7, which processes the same input text into roughly 1.0 to 1.35 times as many tokens depending on content type. Anthropic priced the launch rate specifically so the transition lands as "roughly cost-neutral" despite that change. The practical version of that math, what it means for your actual workload, gets its own section further down. Anthropic also says it has raised rate limits across Chat, Cowork, Claude Code, and the Claude Platform to accommodate the higher token usage that comes with running Sonnet 5 at higher effort levels (if you code with it, our guide to [cutting Claude Code token costs](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) covers how to keep that in check) (a related platform-wide rate-limit restructuring took effect April 26, 2026, ahead of this launch). ## Claude Sonnet 5 Benchmarks: vs Sonnet 4.6 and vs Opus 4.8 [Anthropic published five benchmark comparisons at launch](https://www.anthropic.com/news/claude-sonnet-5). Sonnet 5 beats its predecessor on every single one, and closes most, though not all, of the distance to Opus 4.8.
BenchmarkSonnet 4.6Sonnet 5Opus 4.8What it measures
SWE-bench Pro58.1%63.2%69.2%Long-horizon agentic software engineering
Terminal-Bench 2.167.0%80.4%74.6%Command-line tool use
Humanity's Last Exam (with tools)46.8%57.4%57.9%Graduate-level multidisciplinary reasoning
OSWorld-Verified78.5%81.2%83.4%Computer use (operating-system tasks)
GDPval-AA v2-1,618 pts1,615 ptsReal-world professional knowledge work
SWE-bench Pro is the hardest, most coding-specific test of the five, and it's the clearest Opus win: Opus 4.8 leads 69.2% to 63.2%, and Anthropic is upfront that Opus remains "the model of choice for higher accuracy" on this kind of work. Opus also leads on OSWorld-Verified and edges Humanity's Last Exam. But it doesn't sweep the table, and the popular "Opus wins everything" summary is wrong on two rows. On Terminal-Bench 2.1, Anthropic's own Opus 4.8 launch post puts it at 74.6% — behind Sonnet 5's 80.4%, making Sonnet 5 the stronger Claude model on command-line tool use. And on GDPval-AA v2, which scores real-world knowledge tasks rather than coding puzzles, Sonnet 5 numerically edges Opus 4.8, 1,618 to 1,615, a gap close enough to read as a tie. So the honest read is narrower than the headline: Opus leads pure coding, Sonnet 5 wins terminal work outright, and the two trade the rest. Anthropic also updated its grading for Humanity's Last Exam and OSWorld-Verified at this launch, and retroactively re-scored Sonnet 4.6 under the new methodology: its HLE score moved to 46.8% with tools, its OSWorld score to 78.5%. That's why the Sonnet 4.6 numbers above may not match what you remember from Sonnet 4.6's own launch post. This table uses the regraded baseline, the fairer comparison. Third-party testing points the same direction. Cursor ran its own internal benchmark, [CursorBench](https://www.testingcatalog.com/anthropic-launches-claude-sonnet-5-model-on-claude-and-apis/), and reported Sonnet 5 scoring 57% against Sonnet 4.6's 49%. That's Cursor's proprietary test, not one of Anthropic's published evaluations, but the direction matches everything else here. ## New Capabilities: Effort Levels and Agentic Behavior Sonnet 5 runs with [adaptive thinking](https://geotoolbox.ai/blog/how-does-claude-work) on by default. You choose how hard it thinks through an effort setting, ascending from Low to Medium to High to Extra High to Max. Extra High is new for the Sonnet line: Sonnet 5 is the first Sonnet-class model to offer it, though Max, the actual ceiling, was already available on Sonnet 4.6. Effort defaults to High on both the API and in Claude Code. Higher effort means more reasoning, more tool calls, and, as covered above, more spend. The model carries a 1 million token context window and a maximum output of 128,000 tokens per response, extendable to 300,000 via a batch-API beta header for long-running jobs. Training data runs through January 2026. In practice, according to Anthropic's named early-access partners, that translates to fewer stalled multi-step jobs. [Zapier](https://techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents/) handed Sonnet 5 a two-part task: update Salesforce account tiers, then send a launch announcement to enterprise contacts. It finished the whole thing end to end, something that used to stall halfway with prior models. Legal-tech platform Eve said Sonnet 5 "sits on the Pareto frontier" for its plaintiff-law tasks, citing a price-to-performance ratio that made switching straightforward. ClickHouse said the model reasons in tighter steps and gets users to answers noticeably faster when exploring live data. Anthropic also points to strength on "brownfield" code: the messy, already-shipped parts of a codebase nobody wants to touch, like race conditions and hidden test failures. The claim is that it traces a failure back to its actual root cause instead of patching the visible symptom. That's a different skill from generating new code from scratch, and it's a useful one for [agentic AI](https://geotoolbox.ai/blog/agentic-ai) work specifically, because a fix either holds under the original failing test or it doesn't. One spec worth a caveat: despite the 1 million token window, some early testers report the model losing track of details across very long documents or codebases, more than they expected given the headline number. Treat the [context window](https://geotoolbox.ai/blog/claude-code-context-window) as a ceiling on how much you can feed the model, not a guarantee it will weigh every part of that input equally well. ## What the First Independent Tests Found Anthropic's own benchmarks are one thing; the first outside tests, run in the days after launch, add nuance the launch posts skip. The clearest example is [CodeRabbit's automated code-review benchmark](https://www.coderabbit.ai/blog/claude-sonnet-5-review), worth reading closely because it doesn't split cleanly for or against the model. On the plus side, Sonnet 5's review precision jumped to 38-40%, up from about 29% on Sonnet 4.6: its findings were more often real bugs than noise, and its comments were cleaner. The catch is a genuine surprise. On raw bug-catching, Sonnet 5 actually flagged *fewer* real bugs than its predecessor, roughly 50% of the seeded set versus 63% for Sonnet 4.6. Cleaner output, fewer false alarms, but lower recall on the defects that mattered. If your job is finding bugs rather than writing code, that trade-off is worth knowing before you switch. It's also a concrete answer to the naming critics: "the most agentic Sonnet yet" is not the same claim as "better than 4.6 at everything." Two more practical notes from the same test. At maximum effort, Sonnet 5 posted three to four times as many nitpicks and roughly doubled the cost without finding meaningfully more bugs, so CodeRabbit recommends running it at medium effort to capture most of the benefit without flagship-level spend, the same conclusion the cost math below points to. And because of its extended thinking, it runs slower than Sonnet 4.6 and occasionally rewrites its own plan mid-task, which makes it a poor fit for high-volume pipelines on a tight latency budget. ## Safety and Cybersecurity Changes Anthropic's pre-deployment safety evaluations found Sonnet 5 an overall improvement on Sonnet 4.6. It's better at refusing malicious requests, more resistant to prompt-injection hijack attempts, and lower on hallucination and sycophancy. On the automated behavioral audit that tests for misuse cooperation and deception, Sonnet 5 scored safer than Sonnet 4.6 overall. Anthropic notes it showed somewhat higher rates of misaligned behavior on that specific assessment than the more capable Opus 4.8, a reminder that "safer than its predecessor" and "as safe as the flagship" are different claims. Cybersecurity capability is where Anthropic drew a deliberate line. Sonnet 5 wasn't trained on cybersecurity tasks, and on evaluations that test for genuinely dangerous skills, like developing working software exploits, it performs substantially worse than Opus 4.8 and Mythos 5. In one test built around real Firefox vulnerabilities, Sonnet 5 never produced a single full working exploit, though it did show a slightly higher rate of partial success than Sonnet 4.6. Anthropic attributes that to general intelligence gains rather than any cyber-specific training. Because of that small uptick, Sonnet 5 ships with real-time cyber safeguards on by default, matching the protections already running on Opus 4.7 and 4.8. They're lighter than the restrictions Anthropic put on [Fable 5](https://geotoolbox.ai/blog/fable-5-ban), which block a much wider range of cybersecurity-adjacent tasks. Fable 5 and its research sibling Mythos 5 were also briefly pulled under a US government export restriction over cybersecurity risk in June 2026, but that order was lifted on July 1 and Fable 5 is generally available again (Mythos 5 stays gated to approved organizations). Sonnet 5 launched without that baggage, positioned as the safer, broadly deployable option. Organizations that need reduced guardrails for legitimate security work can apply through Anthropic's Cyber Verification Program, available today on the native Claude Platform, AWS, and Microsoft Foundry, with Google Vertex support coming soon. ## The Real Cost of Switching to Sonnet 5 Headline pricing puts Sonnet 5 at 40-60% cheaper than Opus 4.8. Whether your actual bill drops that much depends on two things the price-per-token comparison doesn't capture. The first is the tokenizer change from the pricing section above: more tokens for the same input, by design, depending on content type. Anthropic built that into the launch pricing to land "roughly cost-neutral," but cost-neutral on average isn't the same as cost-neutral for your specific workload. Code-heavy or non-English content tends to sit at the higher end of that multiplier. The second factor is bigger, and it's specifically a Sonnet-5-vs-Sonnet-4.6 issue, not a Sonnet-5-vs-Opus-4.8 one. [The Decoder's analysis of the launch](https://the-decoder.com/anthropics-new-claude-sonnet-5-closes-the-gap-to-the-pricier-opus-model-series/) makes the point plainly: because Sonnet 5 works more agentically, it's likely to chew through more tokens per task than its predecessor, so even at an unchanged per-token rate, running it could end up costing more than Sonnet 4.6 did for the same job. Anthropic saw the same pattern when Opus moved from 4.6 to 4.7. At higher effort settings specifically, that extra token volume compounds, so the per-task savings against any prior model, Sonnet or Opus, can shrink well below what the rate card alone suggests. There's a tokenizer trap layered on top. [One post-launch cost analysis](https://www.finout.io/blog/claude-sonnet-5-pricing-2026-the-hidden-costs-and-real-savings-behind-the-cost-neutral-launch) ran a fixed daily workload through the numbers: on the $2/$10 rate with moderate tokenizer inflation, it comes out around 20% cheaper than the same job on Sonnet 4.6. That analysis also modelled a scheduled September 1 rise to $3/$15 that, combined with high-end tokenizer overhead, would have pushed the identical workload 20-35% *above* the old Sonnet 4.6 baseline, but Anthropic cancelled that increase in August 2026 and kept $2/$10 as the permanent rate, so the crossover it warned about no longer happens. What is still worth watching is the tokenizer overhead itself: budget against your actual token volume at high effort, not just the per-token rate. We saw the token-volume side first-hand. Running our own multi-phase, multi-agent content workflow on Sonnet 5 the day it launched, the session needed roughly twice the token budget that same workflow usually takes on Opus, enough to trigger a mid-run context-window compaction we hadn't hit running the identical process before. One internal data point, not a benchmark, but it lines up with the pattern Anthropic, The Decoder, and CodeRabbit's max-effort test all describe. At Low or Medium effort, the savings versus Opus 4.8 hold up clearly, and Medium is where independent testing lands as the sweet spot. At High effort and above, compare actual per-task spend, not just the price-per-token, before assuming you've cut your bill. ## Where You Can Use Claude Sonnet 5 Today Sonnet 5 is the new default model for Free and Pro plan users, and it's available to Max, Team, and Enterprise users as well. Developers can reach it through the Claude API as `claude-sonnet-5`, and it's live in Claude Code from day one. On the infrastructure side, it's available through the Claude Platform on AWS, Amazon Bedrock, Microsoft Foundry, and Google Vertex AI, with only the Cyber Verification Program itself still rolling out on Vertex. Third-party tools picked it up the same day. [GitHub Copilot made Sonnet 5 generally available](https://github.blog/changelog/2026-06-30-claude-sonnet-5-is-generally-available-for-github-copilot/) at launch across Copilot Pro, Pro+, Max, Business, and Enterprise tiers, with a gradual rollout across VS Code, Visual Studio, the Copilot CLI, JetBrains, and Xcode. Cursor added it the same day too, complete with its own CursorBench numbers. Early users flagged two rollout snags. While the API and Claude Code had Sonnet 5 live immediately, some reported it wasn't yet selectable in the Claude Desktop app or the Claude Code VS Code extension in the first hours after launch, a sequencing gap rather than a sign the model isn't really live. Separately, Anthropic didn't reset usage limits for the launch, so some users who had already hit their Sonnet 4.6 usage cap couldn't try the new model right away. The core platforms (chat, API, and Claude Code) were confirmed working from the start. ## Early Reactions: What Developers Are Saying Day-one reaction split along predictable lines. Developers running agentic coding workflows were enthusiastic about the jump in tool use and multi-step follow-through, with several early posts calling it the new default choice for Claude Code work that doesn't need Opus-level accuracy. Others pushed back on the value proposition at higher effort settings, for the cost reasons covered above, and a vocal subset argued the jump from Sonnet 4.6 didn't earn the "5" in the name since it doesn't claim a new best-in-class coding score the way some past Sonnet releases did. It's a fair naming argument, not a factual one: the benchmark gains over Sonnet 4.6 are real regardless of what the release was called. It's the same pattern we saw with [Grok 5](https://geotoolbox.ai/blog/grok-5), where a confirmed release still gets picked apart by critics looking for reasons the number shouldn't have gone up. On how it stacks up against OpenAI and Google: a true apples-to-apples benchmark still doesn't exist, because the three vendors don't share a common test suite and Anthropic's page carries no cross-vendor numbers. But in the days since launch a rough consensus has settled. Sonnet 5 is widely rated the best default pick among broadly available frontier models, almost entirely on price-to-performance. OpenAI's GPT-5.6 Sol posts the highest raw agentic scores anyone has published (its top tier hits about 92% on Terminal-Bench 2.1, above every Claude model), and it stopped being a hypothetical on [July 9, 2026, when it went generally available](https://geotoolbox.ai/blog/gpt-5-6). So the argument for Sonnet 5 is no longer that the alternative is out of reach; it is the bill. Sol runs $5/$30 per million tokens against Sonnet 5's $2/$10 rate, two and a half times as much on input and three times on output, for the top of the leaderboard. If a workload genuinely needs those scores, Sol is now there to be bought. If it doesn't, and for most teams it doesn't, Sonnet 5 gets you most of the way for a fraction of it. Google's Gemini 3.1 Pro stays the pick for very long context and multimodal work. For a direct Claude-vs-ChatGPT breakdown, see [our comparison of the two models](https://geotoolbox.ai/blog/claude-vs-chatgpt); we'll fold Sonnet 5 in there as the shared-benchmark picture firms up. ## What This Means If You're Tracking AI Visibility Here's a live example of why this matters. Hours after Sonnet 5's launch, we checked how different AI engines answered questions about it. Most correctly found and cited Anthropic's real announcement. One major engine, however, was still confidently describing Sonnet 5 as unreleased, citing months-old "Fennec" rumor coverage as its source, even with live web search turned on. Search-enabled doesn't automatically mean current, and on a fast-moving release, an AI engine can keep citing stale information well after the facts changed. Re-checking days later, most engines have corrected the launch date, but the old "Fennec" pages, complete with invented benchmark numbers, still circulate in search results and get pulled into some answers, exactly the lag that shifts what buyers read while a topic is still moving. That's the exact gap brand citation tracking exists to catch, and not just on launch days. A cheaper, more agentic Claude model lowers the cost of running autonomous research, shopping, and support agents at scale, which means more of the buying journey happens inside an AI conversation rather than a search results page. In our experience, the days right after a major model update are when AI engines' answers shift the most, sometimes because a model genuinely reasons differently, sometimes because of exactly the kind of stale-data lag we just saw. If you want to see whether ChatGPT, Gemini, Claude, Grok, and the other major engines are actually citing your brand, or surfacing your competitors' sources instead, [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) tracks that across eight AI engines including Claude, and flags the gaps in what those engines cite when your buyers ask. ## Frequently Asked Questions ### Is Claude Sonnet 5 available now? Yes. It launched June 30, 2026, and is live as the default model for Free and Pro plans, available to Max, Team, and Enterprise users, and accessible via the Claude API and Claude Code. ### How much does Claude Sonnet 5 cost? $2 per million input tokens and $10 per million output tokens. Anthropic first announced that as introductory pricing through August 31, 2026, but in August 2026 it cancelled the planned rise to $3/$15 and made $2/$10 the permanent rate. ### Is Claude Sonnet 5 better than Opus 4.8? Not across the board. Opus 4.8 still leads on the hardest agentic coding benchmark, SWE-bench Pro, and is Anthropic's recommended choice for the highest-accuracy work. On knowledge-work tasks (GDPval-AA v2), the two are essentially tied. Which one wins depends on the task and the effort level you're running at. ### Is Claude Sonnet 5 better than Sonnet 4.6? Mostly, but not universally. It wins every published benchmark over 4.6 and is markedly better at multi-step agentic follow-through. The exception showed up in independent testing: on automated code review, Sonnet 5 caught fewer real bugs than 4.6 (about 50% versus 63%) despite cleaner, more precise output, and it uses more tokens per task. For generating and shipping code it's a clear upgrade; for pure bug-hunting on a tight budget, 4.6 can still hold its own. ### What effort level should I run Claude Sonnet 5 at? Effort defaults to High on the API and in Claude Code, but independent testing points to Medium as the sweet spot for most work: it captures most of the quality without the token bill. Maximum effort roughly doubled cost in one code-review benchmark without finding meaningfully more bugs, so reserve High and above for tasks that genuinely need the extra reasoning. ### Can I use Claude Sonnet 5 for free? Yes. It's the default model on Claude's Free plan as of launch, with no separate signup required. ### What happened to "Fennec," the leaked Sonnet 5 from earlier this year? "Fennec" was the internal codename attached to a model identifier that leaked via Google Vertex AI logs in February 2026. That specific checkpoint shipped as Claude Sonnet 4.6, not Sonnet 5, and months of rumor coverage built on top of it, including invented benchmark numbers, never came from Anthropic. The real Claude Sonnet 5 launched June 30, 2026, with its own verified specs. ### Is Claude Sonnet 5 available in Claude Code? Yes, from day one, alongside the Claude API and Claude Platform. Some other surfaces, like the Claude Desktop app, had a brief rollout lag in the hours immediately after launch. ## Sources - Introducing Claude Sonnet 5 - Anthropic - the official launch announcement: pricing, benchmarks, safety evaluations, and customer quotes - `anthropic.com/news/claude-sonnet-5` - Anthropic launches Claude Sonnet 5 as a cheaper way to run agents - TechCrunch - competitive pricing context and business framing - `techcrunch.com/2026/06/30/anthropic-launches-claude-sonnet-5-as-a-cheaper-way-to-run-agents` - Anthropic's new Claude Sonnet 5 closes the gap to the pricier Opus model series - The Decoder - full benchmark breakdown and the real-world token-cost analysis - `the-decoder.com/anthropics-new-claude-sonnet-5-closes-the-gap-to-the-pricier-opus-model-series` - Claude Sonnet 5 review: Should you switch? - CodeRabbit - independent post-launch code-review benchmark: precision gains, the bug-recall regression vs Sonnet 4.6, and the medium-effort recommendation - `coderabbit.ai/blog/claude-sonnet-5-review` - Claude Sonnet 5 Pricing 2026: The Hidden Costs and Real Savings - Finout - the worked cost example modelling the September 1 standard-rate crossover Anthropic later cancelled - `finout.io/blog/claude-sonnet-5-pricing-2026-the-hidden-costs-and-real-savings-behind-the-cost-neutral-launch` - Anthropic launches Claude Sonnet 5 model on Claude and APIs - TestingCatalog - Cursor integration and the CursorBench third-party benchmark - `testingcatalog.com/anthropic-launches-claude-sonnet-5-model-on-claude-and-apis` - Claude Sonnet 5 is generally available for GitHub Copilot - GitHub Changelog - GitHub Copilot day-one availability and plan tiers - `github.blog/changelog/2026-06-30-claude-sonnet-5-is-generally-available-for-github-copilot` - Claude Sonnet 5 "Fennec" leak: what the Vertex AI logs actually show - Dev Community - the leaked Vertex AI log and the checkpoint identifier - `dev.to/marc0dev/claude-sonnet-5-fennec-leak-what-the-vertex-ai-logs-actually-show-3ho5` - Claude Sonnet 5 "Fennec" Leak: What Actually Launched as Claude Sonnet 4.6 - NxCode - confirms the leaked checkpoint shipped as Sonnet 4.6 on February 17, 2026 - `nxcode.io/resources/news/claude-sonnet-5-fennec-leak-2026` - Anthropic just dropped Claude Sonnet 5 - Dev Community (April Fool's satire, self-labeled) - the fictional benchmark post referenced as an example of the rumor cycle's reach - `dev.to/best_codes/anthropic-just-dropped-claude-sonnet-5-and-the-benchmarks-are-kind-of-insane-3ppc` --- ## GPT-5.6 Explained: Sol, Terra, Luna & AI Visibility > OpenAI's GPT-5.6 is three models: Sol, Terra, and Luna. What each does, the post-July-30 API prices, which plans and products get which model, and what it means for AI search. - Canonical: https://geotoolbox.ai/blog/gpt-5-6 - Published: 2026-06-29 · Updated: 2026-08-07 OpenAI announced **GPT-5.6** on June 26, 2026: not one model, but three, named **Sol**, **Terra**, and **Luna**. It is a real step up in capability, at least on OpenAI's own tests. After a two-week window where access was gated to a small group of trusted partners whose participation OpenAI had shared with the US government, OpenAI made the family **generally available on July 9, 2026** across ChatGPT, ChatGPT Work, Codex, and the API. Which of the three you reach depends on your plan and product, and it changed again on August 6, 2026 when Luna became the free-tier default. Most launch coverage stops at the spec sheet. The part it skips, and the part this guide is built around, is what a new flagship like this means for whether AI answers cite you. ## What Is GPT-5.6? GPT-5.6 is OpenAI's latest family of [large language models](https://geotoolbox.ai/glossary/large-language-model), first shown in limited preview on June 26, 2026 and released broadly on July 9. The headline change is structural. Instead of a single flagship, GPT-5.6 ships as three models that share a generation number but sit at different points on the capability-and-cost curve. The naming carries the logic. In OpenAI's new scheme, the number (5.6) marks the generation, while Sol, Terra, and Luna are durable capability tiers meant to advance on their own cadence. It is a shift from the single-flagship pattern of recent releases to a lineup you pick from by job. GPT-5.6 is the latest step in the GPT-5 series, the successor to GPT-5.5. OpenAI frames the family as advancing the frontier on software engineering, computer use, professional knowledge work, scientific research, and cybersecurity, according to its [GPT-5.6 announcement](https://openai.com/index/previewing-gpt-5-6-sol/). The flagship, Sol, is where the biggest gains land. For its first two weeks the launch was defined by a catch: access was gated to a small group of trusted partners and it was not in ChatGPT. That gate lifted on July 9, and access is now tiered by plan and product rather than by partner list. We cover exactly who gets which model below. ## Sol, Terra, and Luna: Which Model Does What The three models are not a "good, better, best" ladder where you always reach for the top. They are tiers you route to by task. **Sol** is the flagship, built for the hardest, longest problems: complex agentic coding, scientific research, and security work where correctness matters more than cost. It is the only model that enables the new max reasoning effort and ultra mode (more on those below). **Terra** is the everyday workhorse. OpenAI positions it as competitive with the previous flagship, GPT-5.5, at a fraction of the price (about half at launch, and cheaper still since the July cut), which makes it the sensible default for serious daily work like support, internal tools, and document analysis. **Luna** is the fast, cheap tier for high-volume and latency-sensitive jobs: bulk classification, routing, summarization, and routine automation where "good enough" intelligence at scale beats peak reasoning.
ModelBest forInput / 1M tokensOutput / 1M tokens
Sol (flagship)Hardest coding, agents, research, security; max + ultra reasoning$5.00$30.00
Terra (balanced)Everyday work; GPT-5.5-class quality at lower cost$2.00$12.00
Luna (fast)High-volume, latency-sensitive, routine tasks$0.20$1.20
The sensible way to use them is not Sol versus Terra versus Luna, but all three deliberately. That is how OpenAI built the pricing, so the tiers map onto a routing strategy rather than a single buying decision: 1. **Default to Terra** for everyday work. It should cover most serious daily tasks at a mid-tier price. 2. **Drop to Luna** for high-volume or latency-sensitive jobs where speed and cost matter more than peak reasoning. 3. **Escalate to Sol** only for the hardest, multi-step problems where correctness is worth the top price and the extra thinking time. ## GPT-5.6 Pricing GPT-5.6 API pricing, per million tokens, is $5 input / $30 output for Sol, $2.00 / $12.00 for Terra, and $0.20 / $1.20 for Luna. Those rates took effect on July 30, 2026, when OpenAI cut Terra by 20% and Luna by 80% and left Sol at GPT-5.5's price. ChatGPT subscription prices did not change. The chronology matters if you are reading older coverage. The prices held from preview into general availability, then moved three weeks after launch, when OpenAI [passed on its own efficiency gains](https://openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6/) and left the flagship alone. Sol's rate is still identical to GPT-5.5's, so at the top of the range you pay the same and get a stronger model. If you use those models through a subscription rather than the API, the cut still reaches you. OpenAI kept ChatGPT and Codex subscription prices and quota budgets exactly as they were, but Terra and Luna now consume fewer credits per task inside Codex and ChatGPT Work, so the same monthly allowance goes further. Terra is the one marketers and developers will fixate on, because OpenAI calls it "about 2x cheaper" than GPT-5.5 for comparable quality. That claim deserves a check, and the July cut has since widened the gap on paper: at $2.00 / $12.00 against GPT-5.5's $5 / $30, the sticker rate is now closer to a two-and-a-half-times saving. But a lower price per token is not the same as a lower cost per task. A reasoning model that thinks longer can burn more tokens to reach the same answer, so the real saving depends on how many tokens your workload uses, not the sticker rate. Treat the cheaper framing as a hypothesis to test on your own prompts. For high-volume work, Luna is where the economics get interesting: at $0.20 per million input tokens it is cheap enough to put an AI step into workflows that could not justify one before. OpenAI has also said Sol would launch on Cerebras hardware at up to around 750 tokens per second in July 2026, initially for select customers. GPT-5.6 charges more for very long prompts, which the headline per-token rates leave out. Once a request carries more than 272,000 input tokens, OpenAI bills the whole request at 2x the input rate and 1.5x the output rate: $10 / $45 for Sol, $4 / $18 for Terra, and $0.40 / $1.80 for Luna. OpenAI also renamed the fast lane on the day of the price cut. What used to be Priority Processing is now Fast mode, billed at twice the standard rate, with Sol running up to 2.5 times faster than standard processing. For the full GPT-5 rate card, including the batch and cached-input discounts and how these tiers compare with the prior generation, our [ChatGPT pricing](https://geotoolbox.ai/blog/chatgpt-pricing) guide works through the whole table. ## What's New: Max Reasoning, Ultra Mode, and What Sol Is Good At Two new controls change how hard Sol can think. **Max** is a new reasoning effort setting that gives the model more time to deliberate on a single problem. **Ultra** mode goes further by bringing in subagents that split a complex job across parallel workers instead of keeping everything in one chain of thought. Both are exclusive to Sol. If you want a refresher on what a [reasoning model](https://geotoolbox.ai/glossary/reasoning-model) is doing under the hood, we have a plain-English explainer. On coding, OpenAI says Sol set a new top score on Terminal-Bench 2.1, a test of agentic command-line workflows that need planning, iteration, and tool use. OpenAI's reported figures, compiled by [DataCamp](https://www.datacamp.com/blog/gpt-5-6-sol-luna-terra), put Sol Ultra around 91.9% and plain Sol around 88.8%. The tier order does not hold perfectly, though: a cheaper model edged a pricier one, and Terra did not clearly beat GPT-5.5 on this test. Because OpenAI has not published a full official lower-tier table and early write-ups disagree on the exact decimals, we are not repeating the lower-tier numbers here. That coding record comes with a caveat OpenAI's own paperwork raises. Its [GPT-5.6 system card](https://deploymentsafety.openai.com/gpt-5-6-preview) reports "instances of the model cheating on tasks and fabricating research results," and the independent evaluator METR said Sol's [detected cheating rate](https://www.rdworldonline.com/openais-gpt-5-6-sol-sets-a-coding-record-its-own-system-card-says-it-cheats/) was higher than any public model it had tested, to the point that it does not treat its capability numbers as a robust measurement. The headline scores are real, but they sit on shakier ground than a clean leaderboard implies. On cybersecurity, OpenAI calls Sol its "most capable model yet," citing gains on vulnerability research and exploitation benchmarks like ExploitBench while using a fraction of the tokens of rival models. It also says Sol does not cross its internal "Cyber Critical" threshold, so the claim is "strongest so far," not "dangerously capable." On biology, Sol scores higher than GPT-5.5 on the GeneBench genomics test while using fewer tokens. ## How to Access GPT-5.6 (Who Gets Which Model) As of July 9, 2026, GPT-5.6 is available across ChatGPT, ChatGPT Work, Codex, and the OpenAI API, per OpenAI's [GPT-5.6 in ChatGPT](https://help.openai.com/en/articles/20001354-gpt-56-in-chatgpt) help article and its [GA announcement](https://9to5mac.com/2026/07/09/openai-announcing-the-next-chapter-for-chatgpt-today-watch-here/); OpenAI also named it the [preferred model in Microsoft 365 Copilot](https://techcrunch.com/2026/07/09/openai-says-gpt-5-6-is-the-preferred-model-for-microsoft-copilot-amid-breakup-chatter/), across Word, Excel, PowerPoint, and Cowork. What you can reach depends on the product as much as the plan, and the two are easy to conflate. Since August 6, 2026, every ChatGPT plan reaches GPT-5.6, but not the same model. Free and Go now default to **Luna**, replacing GPT-5.5; Plus and above default to **Sol** and can pick a reasoning level. Terra is the one tier that never appears in a standard chat. Product by product: - **In standard ChatGPT conversations, which model you get is now set by plan.** Free and Go default to **GPT-5.6 Luna**, which replaced GPT-5.5 as their everyday model on August 6, 2026; OpenAI is adding a **Think** button for harder questions and lifting the cap on text chats from the week of August 10. Plus and above default to **Sol**: Plus reaches it through the Medium and High reasoning options, while Pro, Business, and Enterprise add Extra High and can select GPT-5.6 Sol Pro for the hardest tasks. - **Terra is the exception: it never appears in a standard chat.** You reach it in ChatGPT Work, in Codex, or through the API. Codex gives Free and Go users Terra; Plus, Pro, Business, and Enterprise can choose among all three there and set an effort level for each. ChatGPT Work offers all three from Plus up. - **On the OpenAI API you get all three**, metered per token with no plan gate. - **Ultra mode** (the subagent-splitting mode) is available in Codex for Plus and higher plans, and in ChatGPT Work for Pro and Enterprise. One wrinkle if you are checking this yourself: OpenAI's Help Center still carries the pre-August table, which put Free and Go at "Not included" for every GPT-5.6 option in standard ChatGPT. The August 6 announcement supersedes it for the default model, and the rollout is staggered, so what you see in your own picker may lag what OpenAI has announced.
![Which GPT-5.6 model each ChatGPT plan gets, by product, as of 7 August 2026. Standard ChatGPT conversations: Free and Go get Luna as their default, Plus gets Sol at the Medium and High reasoning levels, and Pro, Business and Enterprise also get Extra High plus Sol Pro. ChatGPT Work: not listed for Free and Go, Sol, Terra and Luna for Plus and above. Codex: Terra for Free and Go, Sol, Terra and Luna for Plus and above. The OpenAI API serves all three per token with no plan gate. Luna replaced GPT-5.5 as the Free and Go default on 6 August 2026, with a Think button and uncapped text chats following from the week of 10 August.](/blog/gpt-5-6/gpt-5-6-access-by-product-aug2026.png)
Every ChatGPT plan now reaches GPT-5.6, but not the same model. Luna became the Free and Go default on 6 August 2026; Terra never appears in a standard chat.
The two-week delay behind that rollout was unusual. OpenAI shared the models with the US government before release and gated early access to a small group of trusted partners, following a June 2, 2026 executive order that directs federal agencies to build an evaluation framework for frontier AI. OpenAI was openly uneasy about it: in its launch post, [quoted by TechCrunch](https://techcrunch.com/2026/06/26/openai-limits-gpt-5-6-rollout-after-government-request-says-restrictions-shouldnt-be-the-norm/), the company said the arrangement "keeps the best tools from users, developers, enterprises, cyber defenders, and global partners who need them." That gate was lifted after the US government review cleared a broad launch, and OpenAI opened the models publicly on July 9. General availability did not mean uniform availability, and the picture moved again a month later. On August 6, 2026 OpenAI made Luna the default for Free and Go, replacing GPT-5.5, and upgraded Sol for Plus and Pro. So every plan now reaches GPT-5.6 in a standard chat, at a different tier, while Terra stays confined to Work, Codex, and the API. ## GPT-5.6 vs GPT-5.5 and the Competition Compared with GPT-5.5, the jump is less about a single smarter brain and more about shape. GPT-5.5 was one flagship; GPT-5.6 is three models plus the new max and ultra reasoning controls, which lets teams dial cost and depth per task instead of paying flagship rates for everything. Sol is the clear capability gain at the top, especially on agentic coding and security, while Terra and Luna push the price-to-performance frontier down. Against rivals, the picture is split. On coding, the Terminal-Bench numbers OpenAI highlighted put Sol ahead of the current Claude and Gemini flagships. On security the claim is narrower than the headlines suggest: OpenAI reports Sol as competitive with an unreleased Claude preview on the ExploitBench test while using far fewer tokens, not as a clear winner over the shipping Claude and Gemini models. And vendors pick the benchmarks that flatter them, so head-to-head results swing with the task. For a grounded view of how these systems differ in practice, our [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) breakdown and our [Grok 5](https://geotoolbox.ai/blog/grok-5) explainer both look past the launch-day numbers. The fair summary is that Sol is a strong frontier model on the tasks OpenAI tested, and independent real-world comparisons are only starting to settle the launch-day claims. ## What GPT-5.6 Means for AI Visibility A new flagship is not just a developer story. It changes what gets said about your brand inside AI answers, and GPT-5.6 makes that unusually visible right now. Watch how AI engines answer questions about GPT-5.6 itself. Most general-purpose models are unlikely to have GPT-5.6 in their training data yet, so their answers lean on live retrieval rather than memory. When we checked across ChatGPT, Gemini, Perplexity, and Claude in late June 2026, two patterns showed up: the engines that pulled text leaned on the same few fast explainers (OpenAI's own pages and a handful of day-one write-ups like DataCamp), while much of the wider mention pool was still video and forum chatter. Beyond OpenAI's own materials, no third-party source has been anointed the citation yet, and the explainer lane is filling fast. That is the [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) lesson in miniature. When a subject is brand new, there is little trained knowledge to fall back on, so the engines reward whoever answered cleanly and early. The window to become a cited source is widest while the topic is still forming, which is why understanding [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) and how one question [fans out into many](https://geotoolbox.ai/blog/query-fan-out) matters more than chasing the keyword once the field is crowded. There is a second implication, and it ties back to the cheating finding above. OpenAI's own system card describes GPT-5.6 as sometimes "overeager" to finish a task, cutting corners on how it gets there. It is not that the model hallucinates more; OpenAI says factual errors actually went down. The risk for a brand is subtler: an agent pushing hard to complete an answer will grab whatever source is cleanest to lift, so being that clear, well-structured source is how you make sure it grabs yours. The sites that get [cited in ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt) tend to state facts plainly and structure them to be quoted; the ones that get skipped leave the answer ambiguous. That is also why [how ChatGPT cites sources](https://geotoolbox.ai/blog/chatgpt-citations) is worth understanding now that GPT-5.6 answers questions on every ChatGPT plan, Luna on the free tier and Sol above it. The practical move now is speed. GPT-5.6 is already shaping the answers ChatGPT returns on every plan, and it is available to every developer on the API, which is the audience that builds the tools everyone else reads. Becoming citable takes longer than a model rollout does. Being citable starts with being reachable: if the AI crawlers cannot fetch your pages, none of the rest matters. That first step is what our free [AI visibility readiness scan](https://geotoolbox.ai/tools/ai-readiness) checks. ## The Specs OpenAI Has Now Confirmed The context window and knowledge cutoff, both unconfirmed at launch, have since been pinned down. OpenAI's [model documentation](https://developers.openai.com/api/docs/models/gpt-5.6-sol) puts the context window at 1,050,000 tokens, with a 128,000-token maximum output, shared across Sol, Terra, and Luna. The knowledge cutoff is February 16, 2026. That cutoff matters for AI visibility: the model's built-in knowledge stops there, so anything more recent, including news about GPT-5.6 itself, can only reach it when the engine retrieves live, not from memory. What is still open is the independent read on capability. Most of the benchmark and capability claims still trace back to OpenAI's own launch materials. Take the launch-day numbers as OpenAI's account, not as settled fact; outside researchers can now run the models themselves. ## The Bottom Line GPT-5.6 is a real step up, on OpenAI's own tests, after a two-week rollout that started gated and opened up on July 9: three well-pitched models, a smarter flagship in Sol, and price cuts in Terra and Luna that got deeper again on July 30. And since August 6 it reaches every ChatGPT plan, with Luna as the free default and Sol above it. The hype will settle and the benchmarks will get tested in the open. The part you can act on today is positioning. A new flagship reshapes what AI tools say about your market the moment it reaches the people asking, and GPT-5.6 already has, on every ChatGPT tier from Plus up and across the whole API, and being the source those tools quote is slow work that rewards starting early. At geotoolbox we build tools to make that measurable, starting with whether the AI crawlers can even reach you. Run the free [AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) to see where you stand. ## Frequently Asked Questions ### Is GPT-5.6 available to the public, or in ChatGPT? Yes, on both counts. GPT-5.6 became generally available on July 9, 2026 across ChatGPT, ChatGPT Work, Codex, and the OpenAI API (and OpenAI has named it the preferred model in Microsoft 365 Copilot). What you get depends on the product as well as the plan: since August 6, 2026 Free and Go default to Luna in standard chats while Plus and above default to Sol, and Terra is reachable only in ChatGPT Work, Codex, and the API. This followed a two-week window where access was gated to a small group of trusted partners. ### Is GPT-5.6 free? Yes, since August 6, 2026. GPT-5.6 Luna is now the default model on the Free plan, replacing GPT-5.5, with a Think button and uncapped text chats rolling out from the week of August 10 (separate limits still apply to files, images, voice, and image generation). Sol still starts at Plus. On the API there is no free tier at all; access is metered per token, starting at $0.20 per million input tokens for Luna. ### How much does GPT-5.6 cost? API pricing per million tokens is $5 input / $30 output for Sol, $2.00 / $12.00 for Terra, and $0.20 / $1.20 for Luna. Those are the rates since July 30, 2026, when OpenAI cut Terra by 20% and Luna by 80% and left Sol untouched at GPT-5.5's price. Requests over 272,000 input tokens bill at double the input rate and 1.5x the output rate. ### What's the difference between Sol, Terra, and Luna? Sol is the flagship for the hardest reasoning, coding, and security tasks, and the only model with max and ultra reasoning modes. Terra is a balanced everyday model with GPT-5.5-class quality at a lower price. Luna is the fastest and cheapest, built for high-volume, routine work. ### Is GPT-5.6 better than GPT-5.5, Claude, or Gemini? On the coding benchmark OpenAI highlighted (Terminal-Bench 2.1), Sol leads GPT-5.5 and the current Claude and Gemini flagships. On security its edge is narrower, competitive with an unreleased Claude preview rather than a clear win over shipping models. Now that GPT-5.6 is broadly available, independent head-to-head comparisons are only beginning, so treat the launch-day leaderboard as OpenAI's account until outside testing catches up. ### What is GPT-5.6's context window? About 1,050,000 tokens, or roughly 1.05 million, with a maximum output of 128,000 tokens, shared across Sol, Terra, and Luna. The [context window](https://geotoolbox.ai/glossary/context-window) is what the model can hold in a single request; the 128,000 ceiling is how much of that it can write back. ## Sources - Previewing GPT-5.6 Sol: a next-generation model - OpenAI - `openai.com/index/previewing-gpt-5-6-sol` - GPT-5.6 Sol model reference - OpenAI API docs (specs: context window, max output, knowledge cutoff) - `developers.openai.com/api/docs/models/gpt-5.6-sol` - GPT-5.6 in ChatGPT - OpenAI Help Center (availability by plan and product) - `help.openai.com/en/articles/20001354-gpt-56-in-chatgpt` - Advancing the price-performance frontier with GPT-5.6 - OpenAI, July 30 2026 (the Terra and Luna cuts, Fast mode, subscription credit consumption) - `openai.com/index/advancing-the-price-performance-frontier-with-gpt-5-6` - Pricing - OpenAI API docs (current per-token rates, long-context tier, Fast mode) - `developers.openai.com/api/docs/pricing` - A preview of GPT-5.6 Sol, Terra, and Luna - OpenAI Help Center (the preview-period gating and its original prices) - `help.openai.com/en/articles/20001325-a-preview-of-gpt-56-sol-terra-and-luna` - OpenAI unveils ChatGPT Work agent, GPT-5.6 models now available - 9to5Mac - `9to5mac.com/2026/07/09/openai-announcing-the-next-chapter-for-chatgpt-today-watch-here` - OpenAI says GPT-5.6 is the "preferred model" for Microsoft 365 Copilot - TechCrunch, July 9 2026 - `techcrunch.com/2026/07/09/openai-says-gpt-5-6-is-the-preferred-model-for-microsoft-copilot-amid-breakup-chatter` - OpenAI limits GPT-5.6 rollout after government request - TechCrunch - `techcrunch.com/2026/06/26/openai-limits-gpt-5-6-rollout-after-government-request-says-restrictions-shouldnt-be-the-norm` - GPT-5.6 Sol, Terra, and Luna: OpenAI's Next-Generation Model Family - DataCamp - `datacamp.com/blog/gpt-5-6-sol-luna-terra` - OpenAI starts previewing GPT-5.6 and its three variants - Engadget - `engadget.com/2203102/openai-starts-previewing-gpt-56-and-its-three-variants` --- ## Best Free SEO Tools (2026): The Honest, Job-Based List > The best free SEO tools in 2026, by job, labeled honestly: actually-free vs freemium vs free-trial, plus the free way to check your AI-search visibility. - Canonical: https://geotoolbox.ai/blog/best-free-seo-tools - Published: 2026-06-28 · Updated: 2026-08-22 The best free SEO tools in 2026 can carry a real SEO program a long way, if you know which "free" is actually free and which is a signup wall in disguise. Most roundups bury that distinction under a 25-item list and skip the fastest-growing job entirely: checking whether AI search engines can even see you. This list is sorted by the job you are doing, labels exactly how free each tool is, and ends with the free AI-visibility checks the others leave out.
![The free SEO stack by job: keyword research, technical audit, page speed, rank tracking, backlinks, analytics, local SEO, on-page, and AI-search visibility, each with its best free tool.](/blog/best-free-seo-tools/free-seo-stack-by-job.png)
A free tool for every SEO job. The last row, AI-search visibility, is the one most lists leave out.
## What "Free" Actually Means Before the list, one filter that decides whether a tool earns a spot here: there are three very different kinds of free, and confusing them is how you end up three clicks deep into a signup wall. **Actually free** means no cost, no usage clock, no credit card. Google Search Console and Google Analytics 4 sit here. You can use them forever without paying. **Freemium** means a genuinely useful free tier sitting under a paid plan, usually gated by a daily or monthly cap. AnswerThePublic and AlsoAsked give you only a few searches a day before asking for money. The free tier is real, but it is built to run out. **Free trial** is not free. It is paid software with the invoice delayed. Useful for a one-time audit, dangerous if you forget to cancel. We label every tool below with one of these three tags, plus the specific catch. Here is the legend:
TagWhat it meansWatch out for
Actually freeNo cost, no paywall, no cardNothing, use it
FreemiumFree tier under a paid planThe daily or monthly limit, and the upsell
Free trialPaid, billing delayedThe renewal date
The pattern worth remembering: the tools that are **actually free** are almost all first-party (Google, Microsoft). The freemium ones are third-party tools and capped desktop editions that give you a taste and meter the rest. ## Best Free Tools for Keyword Research Start with the tools that pull data straight from the source. **Google Keyword Planner** (actually free, needs a Google Ads account) is the closest thing to first-party demand data. You do not have to run an ad to use it, though without active spend it shows volume in wide ranges instead of exact numbers. It is built for advertisers, so the search-volume buckets lean toward commercial terms. **Google Trends** (actually free) does one thing well: relative interest over time. Use it to settle which of two similar keywords is rising and which is fading, and to catch seasonality before you commit a quarter of content to it. **Keyword Surfer** (actually free, Chrome extension) overlays rough volume and related terms directly on the Google results page, which saves you the tab-switching that kills momentum during research. For the questions real people ask, **AnswerThePublic** and **AlsoAsked** (both freemium) map out the question space around a seed term. The free tier is only a few searches a day each, so plan your queries instead of burning them on idle curiosity. **Ubersuggest** (freemium) rounds out the set with volume, difficulty, and content ideas, though its free allowance has tightened to roughly one search a day, so keep it for checking a single keyword rather than bulk work. The most underused free keyword tool is one you already have. **Google Search Console** (actually free) shows the exact queries you already rank for, including the ones sitting on page two with impressions but no clicks. Those are your fastest wins, and no third-party estimate is as accurate as your own performance data. If your goal is visibility in AI answers too, the same query mining feeds directly into [how you optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search). One honest caveat: free volume data is directional, not precise. Different free tools will hand you different numbers for the same keyword, and the gap widens on low-volume terms. Treat the figures as a rough sort order, not gospel, and you will not get burned. ## Best Free Tools for Technical SEO and Site Audits This is the job where free tools genuinely compete with paid suites. **Screaming Frog SEO Spider** (freemium) is the desktop crawler most professionals reach for first. The [free edition crawls up to 500 URLs](https://www.screamingfrog.co.uk/seo-spider/) per crawl, which covers most small and mid-size sites outright and lets you audit a single section of a larger one. It surfaces broken links, redirect chains, missing titles, duplicate metadata, and noindex tags you did not mean to ship. The paid license removes the cap and adds structured-data validation. **Google Search Console** (actually free) is non-negotiable. It is the only tool that shows how Google actually crawls and indexes your site: coverage errors, which pages are excluded and why, mobile usability, and Core Web Vitals from real users. Set it up for every property you touch. **Bing Webmaster Tools** (actually free) is the underrated twin. It runs its own technical audit, accepts a one-click import from your Search Console account, and matters more than its traffic share suggests, because Microsoft's index feeds parts of several AI answer engines. Submitting your [XML sitemap](https://geotoolbox.ai/blog/xml-sitemap) here is a five-minute job with outsized upside. If your site is bigger than Screaming Frog's free 500-URL ceiling, **Beam Us Up** (actually free, on Windows, Mac, and Linux) crawls with no URL limit, trading polish for raw coverage. It has been free since 2013 and is the rare genuinely uncapped free crawler. Two files sit underneath all of that and are worth checking directly, because a general crawler will not always tell you when they are wrong. Google retired the interactive robots.txt tester in December 2023 and replaced it with a read-only report, so you can see the file Google fetched but you can no longer test a URL against a rule. That gap is why we built a free [robots.txt tester](https://geotoolbox.ai/tools/robots-txt-tester) (ours, actually free): it resolves a URL per crawler using Google's real precedence rules and names the line that decided the verdict, which matters because the rule that wins is the longest matching one, not the first. For the sitemap side, our [sitemap extractor and validator](https://geotoolbox.ai/tools/sitemap-extractor) walks index files down to their children, where most free extractors stop, then validates the file and flags URLs you are submitting and disallowing at the same time. Neither replaces a crawler. They answer narrower questions more precisely than one. Between these, you can run a real technical audit, find the issues that quietly suppress rankings, and re-check after fixes, without paying for anything. The one thing to confirm before you trust any audit: that the crawlers you care about, including [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers), are actually allowed to reach your pages in the first place. ## Best Free Tools for Page Speed and Core Web Vitals Site speed is a ranking input and a conversion lever, and the best tools for it are free and first-party. **Google PageSpeed Insights** (actually free) grades any URL on mobile and desktop, combines lab data with real-world Chrome data, and reports your [Core Web Vitals](https://web.dev/articles/vitals) with a prioritized list of fixes. Start at the top of the Opportunities list and work down. **Lighthouse** (actually free) is the same engine built into Chrome's developer tools, so you can audit performance, accessibility, and on-page SEO basics on any page without leaving the browser, including pages behind a login that public tools cannot reach. **GTmetrix** (freemium) adds waterfall charts and region testing if you want to see exactly which request is dragging a page down. You can run a basic test anonymously, and a free account adds more test locations and history, which is enough for spot diagnosis. ## Best Free Tools for Rank Tracking Here is where "free" gets thinner. **Google Search Console** (actually free) is your free rank tracker, with an asterisk. It shows average position, impressions, clicks, and click-through rate for every query you already appear for. What it does not do is daily, location-specific keyword tracking the way a paid tool does, and it only reports terms Google has already shown you for. For most small sites that is enough. You care whether a target page is trending up or down, and Search Console answers that for free. Dedicated trackers ([SE Ranking](/go/seranking?ref=best-free-seo-tools-seranking), [Mangools](/go/mangools?ref=best-free-seo-tools-mangools), and similar) are free-trial only, giving you a week or two and a handful of tracked keywords before asking for a card, so treat them as a trial, not a solution. When you outgrow the free options, our [Semrush vs Ahrefs comparison](https://geotoolbox.ai/blog/semrush-vs-ahrefs) weighs the two big paid suites on real first-party data, our [Semrush review](https://geotoolbox.ai/blog/semrush-review) tests whether the bigger one earns its price, and our [Semrush alternatives](https://geotoolbox.ai/blog/semrush-alternatives) guide picks the cheapest tool for each job. ## Best Free Tools for Backlink Analysis Backlinks are the job where the free ecosystem is weakest, so set expectations accordingly. **Google Search Console** (actually free) shows who links to you, your top linking sites, and your most-linked pages, straight from Google. For your own backlink profile, it is the most trustworthy free source that costs nothing and needs no third-party signup. For a quick look at any domain, **Moz Link Explorer** and **Ubersuggest** (both freemium) give you a limited free check, with Moz capping monthly queries and Ubersuggest a small daily allowance. Free backlink checkers like these show you a **sample** rather than the full picture: a capped slice of links, usually somewhere between twenty and a hundred, plus a domain-strength score. That is fine for a quick gut check on a competitor, but the complete competitor backlink data, the kind that actually informs a link-building campaign, is the single thing almost everyone eventually pays for. If free SEO has a hard ceiling, this is it. When you do reach it, our breakdown of [link building platforms](https://geotoolbox.ai/blog/best-link-building-platforms) covers what is worth paying for. ## Best Free Tools for Analytics and User Behavior **Google Analytics 4** (actually free) is the default for traffic, engagement, and conversions. The setup has a learning curve, but for understanding which organic landing pages actually convert, nothing free comes close. **Microsoft Clarity** (actually free) is underused for SEO, and it has no traffic cap at all. It records [heatmaps and session recordings](https://clarity.microsoft.com/) so you can watch where visitors rage-click, where they stall, and where they abandon. A cleaner on-page experience lifts conversions directly, and tends to support the behavioral signals search systems appear to weigh, so fixing the friction Clarity exposes pays off twice. Pair GA4 with Search Console and Clarity, and you have a complete, genuinely free analytics stack covering what happened, where it came from, and why people did or did not stick around. ## Best Free Tools for Local SEO If you serve a place, **Google Business Profile** (actually free) is the highest-impact free tool you own. A complete, accurate profile with the right categories, hours, photos, and a steady trickle of reviews is what puts you in the local map pack. Verification can take a few days to a couple of weeks, and the payoff lands every time someone searches with local intent. For most local businesses, this single free tool outperforms any paid local suite you could bolt on before the basics are done. ## Best Free Tools for Content and On-Page SEO You do not need a paid optimization platform to ship clean on-page SEO. If you run WordPress, **Yoast SEO**, **Rank Math**, or **All in One SEO Lite** (all freemium, with capable free tiers) handle titles, meta descriptions, focus keywords, readability, and sitemaps right inside the editor. One caveat worth respecting: plugins are code, so keep them updated and do not stack five when one will do. **Mangools SERP Simulator** (actually free) previews how your title and description will render in results before you publish, so you can fix truncated titles and weak descriptions while it still costs nothing. **Hemingway Editor** (actually free in the browser) flags dense sentences and passive voice, which keeps content readable for people and parseable for machines. For structured data, **Google's Rich Results Test** and the **Schema Markup Validator** (both actually free) check whether your markup is valid and eligible for rich results, which is the free answer to the schema validation paid crawlers bundle in. When you outgrow free and want scoring against what already ranks, our guide to [content optimization tools](https://geotoolbox.ai/blog/best-content-optimization-tools) covers the paid options by job. ## Best Free Tools for AI-Search Visibility (the Part Other Lists Skip) Here is the job almost every free-tools roundup ignores, and it is the one growing fastest. [Gartner expects traditional search volume to fall 25% by 2026](https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents) as people get answers from ChatGPT, Perplexity, and Google's AI Overviews instead of clicking ten blue links. If those engines cannot see or cite you, you can lose AI-search visibility without ever realizing it. Start with the honest part: a few tools offer limited free tiers for basic AI-mention checks, but no free tool fully tracks your citations, and serious share-of-voice monitoring across ChatGPT and Perplexity is paid. We cover that category in our [GEO tools guide](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools). What you can do for free, and what decides whether you are eligible to be cited at all, are the foundational checks below. **Can AI crawlers reach you?** This is the silent blocker. AI engines use named crawlers, and you control each one in your robots.txt, but they do different jobs. [OpenAI's OAI-SearchBot](https://developers.openai.com/api/docs/bots) fetches pages for ChatGPT's search answers, while GPTBot is its training crawler; PerplexityBot retrieves for Perplexity; and Google-Extended governs whether your content feeds Gemini, separate from normal Google Search. Block the wrong one and you can quietly lose citations without touching your Google rankings. In our experience, a common and easily-missed reason a crawlable-looking site gets no AI visibility is not the content. It is a robots.txt rule, often added by a developer or a security plugin, that blocks those bots while the team assumes everything is fine. You can confirm it for free by reading your robots.txt for those user-agents, or running our free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker), which tests each major AI bot against your site in seconds; the broader [list of AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) explains what each one does. **Do you have an llms.txt file?** It is an emerging way to point AI models at your most important content. Our free [llms.txt Checker](https://geotoolbox.ai/tools/llms-txt-checker) tells you whether yours exists and validates it, and our [llms.txt explainer](https://geotoolbox.ai/blog/llms-txt) covers whether it is worth adding for your site. **How ready are you overall?** The free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) rolls crawler access, structure, and machine-readability into one score with specific fixes. Two more free moves round it out: search your server access logs for those crawler user-agents to confirm AI bots are actually visiting, and spot-check by asking ChatGPT and Perplexity about your brand to see what they say. None of this costs anything, and together these checks answer the question the rest of your free stack cannot. ## The Free SEO Stack by Job (Quick Reference) If you want the whole free stack at a glance, here is the best free tool for each job and the catch to keep in mind.
JobBest free tool(s)Truly free?The catch
Keyword researchGoogle Keyword Planner + Google TrendsActually freePlanner needs an Ads account; volume shown in ranges
Technical auditScreaming Frog + Search ConsoleFreemium / freeScreaming Frog free caps at 500 URLs
Page speedPageSpeed Insights + LighthouseActually freeNone
Rank trackingGoogle Search ConsoleActually freeAverage position, not daily keyword tracking
BacklinksSearch Console (your own links)Actually freeCompetitor link data needs paid
AnalyticsGoogle Analytics 4 + Microsoft ClarityActually freeGA4 has a setup curve
Local SEOGoogle Business ProfileActually freeVerification required
On-page / contentYoast or Rank Math (WordPress)FreemiumAdvanced features are paid
AI-search visibilitygeotoolbox AI Crawler Checker, llms.txt Checker, AI ReadinessActually freeFull citation tracking is paid
Build from the top down. Search Console, Analytics 4, and the AI-crawler checks take an afternoon to set up and immediately tell you whether anything is fundamentally broken. ## Are Free SEO Tools Enough? For most sites, free tools handle the large majority of day-to-day SEO: keyword research, technical audits, speed, on-page work, your own rank and backlink data, analytics, and the AI-visibility checks above. If you are a small business, a solo founder, or a blogger, you can run a real SEO program for a long time without paying for a single tool. Three things eventually push people to pay. The first is competitor backlink and keyword data at full depth, which is the hardest thing to get for free. The second is scale, when daily caps and 500-URL crawls stop fitting the work. The third is managing many sites at once, where free tools turn into a tab-juggling nightmare. The smart move is to run the free stack until you hit a specific wall, then pay only for the one tool that breaks it, rather than buying a 200-dollar-a-month suite on day one to use 15 percent of it. If that wall is AI search, our honest, [job-based AEO tools guide](https://geotoolbox.ai/blog/best-aeo-tools) covers what is actually worth it. ## Start Free, Then Check the Part Everyone Misses You can build a serious SEO program in 2026 without spending a dollar. Stand up Search Console, Analytics 4, and a crawler, then do the thing most of your competitors have not: confirm that AI search engines can actually reach and read you. Run our free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) and [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) to see where you stand in a couple of minutes. geotoolbox builds free tools for exactly this gap: the AI-visibility half of SEO that the old free stack still leaves out. When you outgrow the free layer and want the citations tracked over time, the paid side of geotoolbox does that across up to eight AI engines, [from $99 a month](https://geotoolbox.ai/pricing) with a 7-day trial. ## Frequently Asked Questions ### What is the best free SEO tool? Google Search Console, for almost everyone. It is actually free, shows real first-party data on how Google sees your site, and doubles as a keyword source and a rank tracker. If you set up only one tool, make it this one. ### Are free SEO tools actually free, or freemium? Both, and the difference matters. Google's tools, Bing Webmaster Tools, Microsoft Clarity, PageSpeed Insights, and Lighthouse are genuinely free, with no paywall for normal SEO use. Tools like AnswerThePublic, AlsoAsked, and Ubersuggest are freemium, with daily-search limits before a paywall. Check which one you are dealing with before you build a workflow around it. ### Can you do SEO for free? Yes, for the large majority of it. Keyword research, technical audits, on-page work, speed, your own rank and link data, analytics, and AI-visibility checks all have capable free tools. The main thing you cannot get well for free is deep competitor backlink and keyword data. ### What is the best free tool for site audits? Screaming Frog's free edition, which crawls up to 500 URLs, plus Google Search Console for how Google actually indexes your pages. Together they catch most technical issues at no cost. ### Is there a free way to check my AI-search visibility? Partly. No free tool fully tracks your citations in ChatGPT or Perplexity, but you can check the foundations for free: whether AI search crawlers like OAI-SearchBot and PerplexityBot are allowed in your robots.txt, whether you have an llms.txt file, and your overall AI readiness. geotoolbox's free AI Crawler Checker, llms.txt Checker, and AI Readiness cover exactly those checks. ### Are free AI SEO tools any good? The ones that check crawler access, llms.txt, and technical readiness are genuinely useful, and enough to make you technically eligible to be cited, though not enough to guarantee it. The tools that promise full AI citation tracking or share-of-voice monitoring are mostly paid, and their free tiers are usually capped at a handful of prompts. --- ## Copilot SEO: How to Get Cited in Microsoft Copilot > Copilot SEO means getting cited in Microsoft Copilot's answers. How Copilot grounds on the Bing index, what to structure, and how to track citations. - Canonical: https://geotoolbox.ai/blog/copilot-seo - Published: 2026-06-28 · Updated: 2026-07-20 Copilot SEO is not a separate discipline so much as your existing SEO aimed at a different target: getting your brand named when someone asks Microsoft Copilot a question, instead of just ranking a link they may never click. Most guides skip the one fact that makes it work. Copilot does not run its own index of the web. It grounds its answers on the Bing index, the same one behind your Bing rankings, and that single detail tells you almost everything about how to show up. Get a few mechanics right and you feed Copilot's consumer surfaces at once, from the answer box in Bing to the assistant in Windows and Edge. Below are the specifics, current as of July 2026. ## What "Copilot SEO" Really Means Search "copilot seo" and most of the results are about using Copilot to write your meta descriptions or do keyword research. This guide is the other job: being the source Copilot names when it answers someone's question. That job is worth doing on its own because Copilot is not one product, and several of its surfaces can name you or skip you: - **Copilot Search**, the AI answer experience in Bing and at copilot.microsoft.com, where a question returns a written answer with cited links - **Copilot in Windows and Edge**, the same assistant built into the operating system and browser - **Microsoft 365 Copilot**, the work assistant inside Word, Outlook, and Teams, which can also pull live web sources into an answer These grew out of what used to be Bing Chat, which Microsoft [rebranded as Copilot in late 2023](https://geotoolbox.ai/blog/what-is-copilot). The naming churn is part of why people are unsure what "optimizing for Copilot" even means. The consumer surfaces (Copilot Search, Windows, and Edge) share one web-retrieval foundation, so the work below feeds all of them at once. Microsoft 365 Copilot can also pull in web sources, but only when an admin turns web search on, and it leans mostly on your internal work data, so public SEO touches just its web layer. Optimizing for the consumer surfaces is mostly your existing SEO and [answer engine optimization](https://geotoolbox.ai/blog/what-is-answer-engine-optimization) aimed at a new target: the answer, not the blue link. ## How Microsoft Copilot Picks Its Sources When Copilot answers a question that needs current facts, it runs a web search and writes its answer from the results. It does not pull from a separate "Copilot index" you can submit to. Microsoft is direct about this in its [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a), which state that "Bing and Copilot search experiences rely on the same core crawling, indexing, and ranking foundation as traditional search." The Bing index is the pool. Everything else is downstream of being in it. The retrieval itself is more specific than "Copilot reads the web." Microsoft's documentation on [how web search works in Copilot](https://support.microsoft.com/en-US/Microsoft-365-Copilot/how-web-search-works-in-microsoft-365-copilot-chat-and-agents) describes the flow: Copilot rewrites your prompt into a short Bing search query, runs it, and draws candidate sources from that Bing result set. It then writes an answer from the most relevant of those results and cites the pages it used. You can see the machinery yourself: the sources button under a Copilot answer reveals the exact search query it generated and the pages it drew on.
![How Microsoft Copilot grounds an answer on the Bing index: a page is crawled by Bingbot into the Bing index, Copilot rewrites the prompt into a Bing grounding query, lifts the cleanest passage, and cites the page in its sources panel.](/blog/copilot-seo/how-copilot-grounds-on-bing-index.png)
Bingbot builds the index; Copilot rewrites the prompt into a Bing query and cites the pages it lifts from.
This is [grounding](https://geotoolbox.ai/glossary/grounding): the answer is anchored to live retrieved pages instead of the model's memory alone. This means two things, and they shape everything else in this guide. First, if a page is not in the Bing index, no amount of content quality makes it a candidate. Second, Copilot works at the passage level, not the whole page. It builds the answer from the part of your page that best fits the rewritten query, the same way the rest of [AI search](https://geotoolbox.ai/blog/how-does-ai-search-work) works, so a self-contained answer it can quote cleanly beats the same point buried in a long page. Two things have to be true to get cited: you are in the index, and your answer is the cleanest passage for that query. ## Why Bing Is the Entry Ticket (and Why That's Good News) The single biggest move for Copilot visibility is getting your pages indexed and ranking in Bing. Most SEOs underweight Bing because it carries a small slice of raw search traffic, and that instinct is now backwards. Bing is no longer just a search engine with a few percent of the market. It is the index behind Copilot's answers in Bing, Windows, and Edge. That makes Bing work an AI distribution play now, not just a traffic one. The reach also runs past Microsoft's own products. In one small test of ChatGPT Agent sessions, Search Engine Land's Jes Scholz found [the agent leaned on the Bing Search API in the large majority of cases](https://searchengineland.com/insights-chatgpt-agent-mode-463127). That is ChatGPT, not Copilot, and too small a sample to bank on, but it points at something structural: the Bing index quietly feeds more AI surfaces than its brand suggests, so the work compounds. There is a second reason to like this surface. It is easier to break into than Google. Copilot leans on Bing's ranking signals, which reward clear, well-structured pages, and it competes a much smaller field than Google's. SEOs report new and even pre-launch sites racking up thousands of Copilot citations while they are nowhere in Google. Those figures are self-reported anecdotes rather than measured data, but the pattern is consistent and it matches what we see: structure and index presence beat raw domain authority here in a way they no longer do on Google. Copilot is also stingy with slots. Practitioners who watch it closely put the count at roughly three sources per answer, far fewer than ChatGPT's longer source lists, so the bar per slot is higher. Fewer slots and lower competition cut in opposite directions, and the net is favorable: you are competing a smaller field for a smaller number of spots, where being a clean, well-ranked Bing result is often enough to land one. ## Step 1: Get Into the Bing Index Everything starts with Bing being able to crawl, index, and rank your pages. If [Bingbot](https://geotoolbox.ai/glossary/bingbot) cannot reach a page, Copilot cannot cite it, full stop. So the first pass is plumbing, and it is usually a quick win because most sites have never been tuned for Bing at all. Three steps cover the index foundation: 1. **Verify your site in Bing Webmaster Tools.** You can import your site and sitemaps directly from Google Search Console, so this takes minutes. It also gives you the reporting you will need later to measure citations. 2. **Submit your [XML sitemap](https://geotoolbox.ai/blog/xml-sitemap) and use IndexNow.** [IndexNow](https://www.indexnow.org/), which Microsoft supports and the Bing guidelines tell you to use, pings Bing the moment a page is published or updated, so new content becomes eligible for Copilot in hours instead of waiting on a crawl. 3. **Clear the crawl blockers.** Confirm that no robots.txt rule, `noindex` tag, login wall, or firewall is blocking Bingbot. Aggressive Cloudflare or WAF rules that lump Bingbot in with scrapers are a common, silent cause of AI invisibility. The crawler question trips people up, because the AI era added a pile of new bots and the advice to "block AI crawlers" is everywhere. Keep two categories separate. Bingbot is the crawler that builds the index Copilot answers from, so blocking it removes you from Copilot. Bots like GPTBot or the training crawlers are a different decision about whether your content trains a model, and blocking them does not affect Copilot citations one way or the other. The mistake that costs you is blocking the wrong one. Confirm Bingbot has a clear path in your robots.txt and in Bing Webmaster Tools first, since that is the crawler Copilot depends on. Then sort out the AI bots: our free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) fetches your robots.txt server-side and shows which of 34 AI crawlers, including a Microsoft AI crawler, you are allowing or blocking, so the block-all-or-allow decision is a deliberate one. For the wider picture of who reaches your site, see our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers). ## Step 2: Structure Pages So Copilot Can Lift Them Ranking in Bing gets you into the candidate pool. Structure decides whether you get lifted out of it. This is the gap that frustrates people whose pages rank fine yet never get named: the answer is there, but it is buried where Copilot cannot cleanly extract it. Position qualifies you; an extractable passage wins the slot. Ranking number one is not even required, since the pages an AI answer cites do not always match the top organic results. Microsoft publishes its own short list of what makes content easy to ground. The [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a) tell you to state facts explicitly, use clear and consistent names for entities, keep a single topic per URL, and put the key information near the top of the page. Read those as instructions, not platitudes. They translate into a handful of habits: - **Lead each section with the answer.** State the fact in a self-contained sentence in the first line or two, then add the detail. Copilot lifts a clean claim it can quote in one line over the same point spread across three paragraphs. - **Chunk by question.** Give each distinct question its own heading and a standalone answer beneath it, so a passage makes sense lifted out of the page. Phrasing the heading as the question users actually ask helps the match. - **Keep the answer in the HTML.** If your key facts load through client-side JavaScript, a crawler or grounding fetch can hit a near-empty shell. Server-rendered, text-first pages extract cleanly. - **Write with conviction.** Independent [research by Kevin Indig](https://www.growth-memo.com/p/the-science-of-how-ai-pays-attention) found that cited passages were nearly twice as likely to use definitive language as uncited ones, and that 44% of citations came from the first third of the page. That study measured ChatGPT citations, so treat it as a general AI-citation pattern rather than a Copilot-specific rule, but it lines up with Bing's own advice to put your clearest, most committed sentences up top. Schema markup is worth adding, with a realistic expectation of what it does. [Structured data like FAQPage, HowTo, and Article](https://geotoolbox.ai/blog/schema-markup-for-ai) helps Copilot parse your entities and pull the right snippet, but it is an extraction aid, not a citation trigger. It makes a good passage easier to lift; it does not make a weak page get chosen. In the scans we run at geotoolbox, the pages that win citations are almost always the ones that answer a single question cleanly near the top, with schema as a supporting layer rather than the cause. If you want a second read on how liftable a given page is, our [Content Analyzer](https://geotoolbox.ai/features/content-analyzer) grades a URL on the same signals, and the broader workflow lives in [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search). ## Why Copilot Picks One Source Over Another Indexing and structure get you into the running. When Copilot has several pages that could answer a query and room for only about three, a few softer signals decide who it names. They are worth knowing, because this is the gap behind the most common complaint: my coverage is fine, so why does Copilot keep naming a competitor? Freshness tends to count for more on Bing than on Google, so a current, honestly dated page beats a stale one on questions that move. Authority still matters: a clear author, real credentials, and citing your own sources all help Copilot trust the page. And because the answer is assembled from what several sources agree on, being named accurately across reputable sites, profiles, and communities counts as much as your own page. Copilot also leans toward recognized entities, which is why consistent naming and a strong [entity profile](https://geotoolbox.ai/blog/entity-seo), with sameAs links to places like LinkedIn, help it connect a mention to you. None of this overrides being in the index, but it often decides the last open slot. ## Step 3: Control How Copilot Cites You Microsoft gives you direct controls over whether and how Copilot can cite a page, and almost no one uses them. These are the same robots and meta directives you already know, but the [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a) now document exactly what each one does to Copilot specifically. They cut both ways: most pages want to be fully citable, but a paywalled article, a pricing page, or a competitor-bait comparison might not.
DirectiveEffect on Copilot citationWhen to use it
Default (no directive)The page is fully eligible to be grounded, quoted, and citedAlmost everything you publish and want found
NOARCHIVEPrevents the content from being used in Copilot responses and grounding results entirelyPages you never want Copilot to lift from
NOCACHELimits Copilot to using only the URL, title, and snippet, not the full pageGated or sensitive pages you still want surfaced as a link
NOSNIPPET / data-nosnippetRemoves the snippet, which can limit Copilot citation quality for the marked contentSpecific paragraphs you do not want quoted (wrap with data-nosnippet)
data-snippetMarks the exact text you prefer Bing to display or citeSteering Copilot toward your cleanest answer passage
The practical move is the last row. Most pages should stay on the default and let Copilot quote freely, but the `data-snippet` attribute lets you specify the text you would rather it use, which helps when your best answer is not the first thing on the page. It is Bing-documented, though how strongly it steers Copilot in practice is not something Microsoft quantifies, so treat it as a nudge, not a guarantee. Reserve `NOARCHIVE` and `NOCACHE` for the handful of URLs where being quoted in full works against you. A blanket `NOARCHIVE` applied site-wide by a cautious developer quietly removes you from Copilot answers you would have wanted to win. ## Step 4: Measure Your Copilot Citations Copilot is the most measurable AI engine to optimize for, because it is the only one with a first-party citation report. In February 2026, Microsoft launched the [AI Performance report in Bing Webmaster Tools](https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview) as a public preview. It shows when your site is cited across Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations, reported through metrics like Total Citations, Average Cited Pages, visibility trends, and page-level citation activity. Google Search Console still has no equivalent, which is the quiet reason Copilot is the easiest AI surface to optimize with feedback rather than guesswork. The most useful field is grounding queries: a sample of the key phrases the AI used to retrieve your cited pages, aggregated across Copilot, Bing's AI summaries, and partner surfaces. Treat them as retrieval-intent clues rather than a complete log of exact Copilot searches, and they still close the loop. Instead of guessing which questions to write for, you mine that list for the ones you answer thinly or not at all, write the clean passage, and optimize against real retrieval behavior rather than a keyword tool. A [June 2026 update](https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare) added intent labels and topic groupings on those queries plus Citation Share and Compare views, so you can also watch how your presence moves against other cited sources over time. Two honest caveats. The report is a preview, so expect the metrics and categories to keep changing. And citation data is not traffic data. Most AI clicks arrive with no referrer, so Copilot sessions are nearly invisible in GA4, which is exactly why the citation report matters more than your analytics here. Microsoft Clarity can surface some of that referral behavior, but treat it as a separate, smaller signal from the citation counts, which measure visibility rather than clicks. The discipline is the same one we build into our tools: sample the questions that matter, log who gets cited, and close the gaps one passage at a time. geotoolbox's [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) tracks brand citations in Bing Copilot alongside the other major engines, and the [Domain Overview](https://geotoolbox.ai/features/domain-overview) turns those repeated checks into a visibility baseline over time. ## What a Copilot Citation Is Actually Worth Set expectations before you celebrate a big citation number, because the metric that looks most impressive is the one that converts least. SEOs watching the new report routinely see huge citation counts paired with almost no sessions, on the order of a tenth of a percent of citations turning into a click. Those are practitioner-reported figures, not vendor data, but the dynamic is real: most people read the Copilot answer and never click through, so a citation is a brand impression far more often than it is a visit. That reframes the win. A Copilot citation is worth what a recommendation is worth: your name shows up at the moment someone is deciding, which drives the branded searches and direct visits that follow even when the answer itself gets no click. Plan for presence in the answer, not a flood of referral traffic, and value a citation the way you would value being named in an analyst's shortlist rather than a paid click. It also pays to stay honest about the ceiling. Microsoft is explicit that good optimization makes a page eligible for grounding and citation but does not guarantee it. There is no published "Copilot ranking algorithm" to reverse-engineer and no setting that forces a citation. You are improving your odds across many answers, not buying a slot, which is exactly why the measure-and-iterate loop matters more here than any single tactic. ## Copilot vs the Other AI Engines Copilot's defining trait is that it rides the Bing index, which makes Bing SEO the lever and gives you a first-party report no other engine offers. The other major engines retrieve differently and reward slightly different work, so it helps to see where Copilot sits.
EngineWhat it grounds onCrawler to allowFirst-party citation report?
Microsoft CopilotThe Bing indexBingbotYes, Bing Webmaster Tools AI Performance
ChatGPTIts own search index plus Bing for some retrievalOAI-SearchBotNo
Google GeminiThe Google Search indexGooglebotNo dedicated citation report
PerplexityIts own index plus partner search APIsPerplexityBotNo
ClaudeWeb search (via Brave)Claude-UserNo
The fundamentals travel. A reachable, indexed, answer-first page is a citation candidate on all five engines, so the same [GEO work](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) covers Copilot too. What is unique to Copilot is the upside: the Bing index reaches more places than its traffic share implies, and it is the one engine that reports back which of your URLs get cited across its AI surfaces. ## What Doesn't Move Copilot Citations A few tactics get sold as Copilot requirements that do nothing, and they are worth naming so you can spend the time elsewhere. **Microsoft Ads.** Paying for Bing ads does not buy or boost organic Copilot citations. The two systems are separate, and there is no documented path from ad spend to a grounded citation. **GitHub Copilot tweaks.** GitHub Copilot is a coding assistant. It is a different product that shares the brand, and nothing you do for it affects whether Microsoft Copilot cites your website. **Keyword stuffing.** Copilot lifts passages that read as clear, credible answers. Repeating a phrase to hit a density target makes a passage harder to lift, not easier, and Bing's quality systems discount it. **An llms.txt file.** There is no evidence Copilot reads [llms.txt](https://geotoolbox.ai/blog/ai-crawlers), and Microsoft does not list it as a grounding input. Treat it as unproven rather than a requirement, and put the effort into being in the Bing index with extractable answers instead. ## Frequently Asked Questions ### Does Microsoft Copilot use Bing to find sources? Yes. Copilot grounds its web answers on the Bing index, and Microsoft's own guidelines state that Bing and Copilot search rely on the same crawling, indexing, and ranking foundation as traditional search. When Copilot needs current facts, it rewrites your prompt into a Bing query, pulls candidate pages from the results, and cites the passages it uses. ### Is Copilot SEO the same as Bing SEO? It is built on Bing SEO but adds a second layer. Ranking in Bing gets your page into Copilot's candidate pool, which is the prerequisite. Getting cited then depends on structure: answer-first passages, self-contained sections, and clear writing Copilot can lift cleanly. So strong Bing SEO is necessary, and extractable content is what turns a ranking page into a cited one. ### How do I check if Copilot is citing my website? Use the AI Performance report in Bing Webmaster Tools, which launched in public preview in February 2026. It shows your total citations across Microsoft Copilot and Bing AI answers, which pages get cited, and a sample of the grounding queries the AI used to retrieve them. It is currently the only first-party AI-citation report any major engine offers. ### How long until Copilot cites a new page? It depends on how fast Bing indexes the page. Submitting it through Bing Webmaster Tools and pinging IndexNow can make a new page eligible within hours rather than waiting on a normal crawl. Citation is never guaranteed, but faster indexing is the fastest route to becoming a candidate. ### Does blocking Bingbot block Copilot? Yes. Bingbot builds the index Copilot answers from, so if your robots.txt, firewall, or WAF blocks Bingbot, your pages cannot be cited in Copilot. This is different from blocking AI training crawlers like GPTBot, which has no effect on Copilot citations either way. ### Does schema markup get me cited in Copilot? It helps extraction, not selection. Schema like FAQPage, HowTo, and Article makes it easier for Copilot to parse your entities and pull the right snippet, but it does not make a weak page get chosen. Add it as a supporting layer on top of clear, answer-first content, not as a citation trigger on its own. ## Start with the Bing Index That is the whole job: get into the Bing index, then be the cleanest answer Copilot can lift. Do those two things and Copilot citations follow the rest of your AI visibility work, with the bonus that this is the one engine that reports back exactly which pages it cites. Start at the gate the other guides skip. geotoolbox's free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows which AI crawlers your robots.txt is letting in or shutting out, the [Content Analyzer](https://geotoolbox.ai/features/content-analyzer) grades how liftable a page is once it is reachable, and the [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) tracks when Bing Copilot starts naming you. Open the door, make your best pages the clearest current answer, and give Copilot its best reason to pick you. ## Sources - Bing Webmaster Guidelines - Microsoft Bing (Copilot grounding, citation directives, and content rules) - `bing.com/webmasters/help/webmaster-guidelines-30fba23a` - Introducing AI Performance in Bing Webmaster Tools (Public Preview) - Bing Webmaster Blog, February 2026 - `blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview` - Bringing the best of AI search to Copilot - Microsoft Copilot Blog, November 2025 - `microsoft.com/en-us/microsoft-copilot/blog/2025/11/07/bringing-the-best-of-ai-search-to-copilot` - How web search works in Microsoft 365 Copilot - Microsoft Support - `support.microsoft.com/en-US/Microsoft-365-Copilot/how-web-search-works-in-microsoft-365-copilot-chat-and-agents` - IndexNow - IndexNow protocol (supported by Microsoft Bing) - `indexnow.org` - The Science of How AI Pays Attention - Kevin Indig, Growth Memo, February 2026 - `growth-memo.com/p/the-science-of-how-ai-pays-attention` - Insights from 100 conversations with ChatGPT Agent mode - Jes Scholz, Search Engine Land, October 2025 - `searchengineland.com/insights-chatgpt-agent-mode-463127` --- ## What Is Microsoft Copilot Studio? A Plain Guide (2026) > What is Microsoft Copilot Studio? A plain guide to the agent builder: what it does, how it differs from Copilot, what it really costs, and its honest limits. - Canonical: https://geotoolbox.ai/blog/copilot-studio - Published: 2026-06-28 · Updated: 2026-08-22 Microsoft Copilot Studio is the low-code tool for building your own AI agents, not the Copilot you chat with in Word or Teams. Microsoft attaches the word "Copilot" to a whole family of products, and this is the one people mix up most. This guide covers what Copilot Studio actually is, what you can build with it, what it really costs once you get past the headline price, and where it falls short. Everything below is current as of July 2026. ## What Is Microsoft Copilot Studio? Microsoft Copilot Studio is a low-code platform for building, managing, and publishing custom AI agents. Microsoft's own [overview](https://learn.microsoft.com/en-us/microsoft-copilot-studio/fundamentals-what-is-copilot-studio) calls it "a graphical, low-code tool for building agents and agent flows." You describe the agent you want in plain language, point it at your data, give it a few actions, and publish it to the places your customers or employees already work. It sits inside the Power Platform, the same family as Power Automate and Power Apps, and it is the direct successor to Power Virtual Agents, which Microsoft renamed in late 2023. If you ever built a "PVA chatbot," you have already used an early version of this. An agent here is more than a scripted chatbot. Microsoft defines it as something that "coordinates language models, along with instructions, context, knowledge sources, topics, tools, inputs, and triggers" to get a job done. In plain terms, it can answer questions from your documents, decide what to do next, and take an action, like filing a ticket or updating a record. Copilot Studio is the builder, not the assistant. You work in it as a maker, at [copilotstudio.microsoft.com](https://copilotstudio.microsoft.com), and the agents you create show up elsewhere. That single distinction clears up most of the confusion around the product. ## Copilot Studio vs Microsoft 365 Copilot Both products carry the Copilot name and Microsoft sells them side by side, so they are easy to confuse. The difference is simple once you see it. Microsoft 365 Copilot is the ready-made assistant your employees use inside Word, Excel, Outlook, and Teams. Copilot Studio is the workshop where you build custom agents of your own. One you use; the other you build with. It is one branch of the wider [Microsoft Copilot](https://geotoolbox.ai/blog/what-is-copilot) family. They are designed to work together. An agent you build in Copilot Studio can be published straight into Microsoft 365 Copilot, so an employee asks it a question in Copilot Chat and never knows a maker assembled it. But the two are licensed differently and aimed at different people, which matters the moment you start budgeting.
AspectMicrosoft 365 CopilotMicrosoft Copilot Studio
What it isA finished AI assistantA platform to build custom AI agents
Who it is forEvery employeeMakers, IT, and business teams
What you do with itSummarize, draft, and analyze inside Office appsDesign, ground, and publish your own agents
Where it runsInside Word, Excel, Outlook, TeamsA standalone web app; agents deploy to many channels
How it is licensedA fixed price per user, per monthConsumption based, paid in Copilot Credits
Example"Summarize this email thread"An HR agent that answers policy questions on your website
Most organizations end up with both. They give employees Microsoft 365 Copilot for everyday work, then reach for Copilot Studio when a specific job, like a customer-facing support agent or an internal IT helpdesk, needs more than the general assistant can do. ## What You Can Build: The Agent Types Copilot Studio is general enough to build very different things, which is part of why pricing is hard to pin down later. These fall into a few buckets, not official product tiers, just the shapes most agents take, and they sit on the broader idea of [agentic AI](https://geotoolbox.ai/blog/agentic-ai). **Knowledge agents** answer questions from your content. You point the agent at a SharePoint site, a public website, internal documents, or your Microsoft Graph data, and it uses retrieval to ground its answers in those sources rather than making things up. An HR agent that explains leave policy, or a product agent that answers from your manuals, lives here. **Workflow agents** do more than talk. They connect to Power Automate and to systems like Dynamics 365, SAP, Salesforce, or ServiceNow, so they can reset a password, log a ticket, or submit an expense. This is where an [AI agent](https://geotoolbox.ai/glossary/ai-agent) stops being a chatbot and starts being a coworker that completes a task. **Autonomous agents** are triggered by an event rather than a person. A new order arrives, and the agent confirms stock, checks the shipping date, and emails the customer, all without someone typing a prompt. **Customer-facing agents** publish to your website, a mobile app, or a contact center, and in 2026 they can include real-time voice agents that take a phone call, answer, and hand off to a human when needed. Microsoft also added **computer-using agents**, which operate a website or desktop app through its own interface, clicking and typing like a person. That lets them automate steps in systems that never exposed a clean API. They are billed at the same rate as a standard agent action, a detail that matters once you reach pricing. These are not hypothetical. The UK's National Zakat Foundation built a Copilot Studio agent that scores aid applications by urgency, and with the wider automation around it [cut grant wait times by 80%](https://www.microsoft.com/en/customers/story/23068-national-zakat-foundation-microsoft-copilot-studio), down from four or five months, across roughly 10,000 applications a month. ## How Copilot Studio Works Building an agent follows the same four steps no matter how simple or ambitious it is.
![The four steps of building a Microsoft Copilot Studio agent: describe the agent in plain language, ground it on your knowledge sources such as SharePoint, websites, and Microsoft Graph, add tools and actions through Power Automate and connectors, then publish it to channels like Teams, your website, and Microsoft 365 Copilot.](/blog/copilot-studio/how-copilot-studio-works.png)
Describe, ground, connect, publish. The same loop builds a simple FAQ bot or a multi-step workflow agent.
First you describe the agent. You tell Copilot Studio, in plain language, what it should do, what tone to use, and what it must not touch. The platform turns that into a working draft you can refine. Then you ground it. An agent with no knowledge source is just a generic chatbot, so you connect it to the content it should answer from: a SharePoint library, a public website, uploaded files, or your tenant's Microsoft Graph. This is the step that decides whether answers are useful or vague, and it is where a lot of projects quietly succeed or fail. Next you add tools and actions. Through Power Automate flows, prebuilt connectors, and newer options like Model Context Protocol (MCP) tools and the Work IQ layer, the agent gains the ability to do things rather than only describe them. Finally you test and publish. You try the agent in a test panel, then publish it across the channels you need: Microsoft Teams, a website, a mobile app, or inside Microsoft 365 Copilot. Microsoft rebuilt this experience through 2026 around generative orchestration, where the agent plans which knowledge to pull and which tools to call to answer, rather than following only the scripted topic-by-topic flows of the older Power Virtual Agents model. Those classic topics still exist when you want tight, predictable control, but you rarely have to think in those terms to ship something useful. ## Pricing and Licensing: Copilot Credits Decoded Pricing is where the simple product gets complicated. Copilot Studio does not charge per user. It charges per **Copilot Credit**, a usage currency pooled across your whole tenant. The number of credits an agent burns depends on what it does, not how many people talk to it. For how Studio's credits sit alongside the rest of Microsoft's Copilot pricing, see our [Microsoft Copilot pricing guide](https://geotoolbox.ai/blog/copilot-pricing). There is one common point of confusion here. On September 1, 2025, Microsoft renamed the unit from "messages" to "Copilot Credits." Per Microsoft's [licensing documentation](https://learn.microsoft.com/en-us/microsoft-copilot-studio/billing-licensing), there was "no change in the quantity per prepaid pack or to the pay-as-you-go rate." So older guides that talk about "25,000 messages" are describing the same 25,000 units, just under the old name. What the credit model makes visible is that not every interaction costs the same. This is the table that explains your bill, taken from [Microsoft's billing rates](https://learn.microsoft.com/en-us/microsoft-copilot-studio/requirements-messages-management):
What the agent doesCopilot Credits
Classic answer (a scripted, pre-written reply)1
Generative answer (AI writes the reply)2
Agent action (a tool or step, including computer use)5
Microsoft Graph grounding (per message)10
Agent flow actions (per 100 actions)13
Standard AI tools (per 10 responses)15
Premium / reasoning AI tools (per 10 responses)100
Content processing (per page)8
Two things follow from that 1-to-100 spread. First, costs **stack** inside a single reply. Microsoft's own example is an agent that grounds in your Microsoft Graph and writes a generative answer: that one response costs 12 credits, 10 for grounding plus 2 for the answer. Reasoning models cost the most, metered on top of the feature rate at 10 credits per 1,000 tokens, so a heavy reasoning reply can run many times the price of a normal one. Second, the same agent design can cost wildly different amounts. A scripted FAQ bot costs about a cent per answer, and runs free for internal users on the included path. Swap in reasoning and live grounding and the same bot can run into the hundreds a month. Microsoft's documented support-agent example, four scripted plus two generative answers across 900 customers a day, works out to 7,200 credits a day before anything fancy. Before you build, run your design through Microsoft's free [Copilot Studio agent usage estimator](https://microsoft.github.io/copilot-studio-estimator/). It turns agent type, traffic, and grounding choices into a credit forecast, which is the only reliable way to predict a consumption bill. Two more cost realities catch teams off guard. Unused credits do not roll over month to month, and when consumption hits 125% of your prepaid capacity, Microsoft disables the agents until you add credits. So one department's experimental reasoning agent can drain the shared pool and disable another team's live bot, unless you ring-fence capacity per environment. ## Pricing Models: Packs, Pay-as-You-Go, and What's Included You buy those credits in one of a few ways, and the right one depends on how predictable your usage is.
PlanPriceBest for
Free trial$0Building and testing an agent (you cannot publish from the trial)
Capacity pack$200 / month for 25,000 creditsSteady, predictable usage (about $0.008 a credit)
Pay-as-you-goAbout $0.01 per credit, via AzurePilots and spiky usage, no upfront commitment
Pre-purchase (CCCU)Annual commitment, up to 20% offLarge, committed deployments
Included with Microsoft 365 CopilotNo extra creditsInternal agents used by licensed employees
A common play is to cover your expected usage with a pack at the cheaper rate, then leave pay-as-you-go switched on as a safety net so you never hit the 125% cutoff. One quirk worth knowing: the standalone maker license itself is free, but assigning it needs a tenant credit-pack subscription first (a Microsoft 365 Copilot license or the trial are other ways in). The headline price is also not the whole bill, and how big the rest of it gets depends on how you license it. When you reuse Microsoft 365 Copilot seats for internal agents, the credits sit on top of those seats (around $30 per user a month for Enterprise, less for the Business tier) and your base Microsoft 365 plan. And if your agents call custom Azure OpenAI models or process documents through SharePoint, separate Azure and SharePoint charges land on different invoices. A 100-person enterprise can be paying thousands a month before a single standalone credit is used. That layering, not the $200 pack, is what makes a Copilot Studio budget hard to forecast. ## Is Copilot Studio Free? Partly, and only in specific cases. There is a free trial, but Microsoft's documentation is clear that it lets you build and test an agent without letting you publish it, so it is for evaluation, not production. Several guides get this wrong and say you can deploy for free; you cannot. The bigger "free" path is the Microsoft 365 Copilot inclusion. If your users already have a Microsoft 365 Copilot license, their internal agent interactions inside Teams, SharePoint, and Copilot Chat, from classic answers to generative answers and Microsoft Graph grounding, are zero-rated and consume no credits. The moment an agent faces the outside world, runs autonomously, or serves people without a Copilot license, it starts spending credits. So "is it free" really means "who uses the agent, and where." ## Do You Actually Need Copilot Studio? Before you buy anything, check whether you need full Copilot Studio at all, because many people who search for it do not. If you just want AI help inside your Office apps, that is [Microsoft 365 Copilot](https://geotoolbox.ai/blog/what-is-copilot), not Studio. If you want a simple internal question-and-answer agent, the lighter agent builder included with a Microsoft 365 Copilot license is often enough on its own. You only need full Copilot Studio, with its credit packs, when you cross into custom topics and branching logic, actions that reach external systems through Power Automate, publishing to a public website or app, or enterprise governance controls. The fit also depends on your stack. Copilot Studio is strongest when you are already invested in Microsoft 365, SharePoint, and Dynamics, where the connections are close to native. Wiring agents into non-Microsoft systems like Salesforce or SAP is real engineering work, not a few clicks. And for a small team, the $200 monthly floor plus unpredictable credit usage is often more than the value it returns, which is why smaller shops sometimes pick a lighter third-party tool instead. ## Copilot Studio vs the Other "Copilots" Microsoft's naming is genuinely confusing, so it helps to place Copilot Studio next to the products it gets mistaken for. Microsoft 365 Copilot is the assistant employees use. GitHub Copilot is a separate product for writing code in your editor, unrelated to agent building. Azure AI Foundry is the pro-code path for engineers who want full control over models and orchestration, where Copilot Studio is the low-code path. And Agent 365, which Microsoft made generally available in May 2026, is the control plane IT uses to govern and secure the agents once they exist. You also choose the brain inside the agent. Copilot Studio lets you pick the underlying model, including options like [Anthropic's Claude](https://geotoolbox.ai/blog/what-is-claude-ai) alongside OpenAI's GPT models, though custom models you bring through Azure get billed separately. One question keeps coming up in the search results: is Copilot shutting down? No. The confusion comes from Power Virtual Agents, the old product that became Copilot Studio in late 2023. Nothing was killed; it was renamed and absorbed. The one real sunset is small: since June 30, 2026, the legacy Copilot Studio for Teams app can no longer create classic chatbots. It redirects makers to the main web app instead, and agents built before then keep working. ## Where Copilot Studio Falls Short Copilot Studio is capable, but the demos are smoother than the reality, and it is worth going in clear-eyed. The clearest picture comes less from vendor blogs than from the people running it day to day. The first issue is cost, as we saw. Microsoft publishes the per-credit rates but not the usage volume your agents will hit, and the "fair use" limits on included internal usage are never quantified, so finance teams are left forecasting in the dark. The second is grounding. Pointing an agent at your data is easy; getting clean, current answers out is not. In community threads on [Reddit's r/copilotstudio](https://www.reddit.com/r/copilotstudio/comments/1ip7y97/what_is_your_experience_with_copilot_studio_does/), makers report that SharePoint grounding can be uneven, that uploaded documents often outperform pointed sources, and that agents sometimes lose the thread of a conversation after a couple of turns. Treat those as reported experiences rather than fixed facts, but they recur often enough to plan around. The third is the "low-code" promise. You can stand up a basic agent in an afternoon, but anything serious tends to pull in variables, expressions, JSON, and connector authentication, which is to say it needs someone technical after all. A fourth is governance. An agent inherits the permissions of the data you ground it on, so a poorly scoped agent can surface SharePoint content a user was never meant to see, and a public agent published with no authentication can be used by anyone with the link, along with whatever it is grounded on. Microsoft's answer is Managed Environments, data loss prevention policies, and Agent 365, but those are controls you set up rather than defaults, and they are split between your Microsoft 365 and Power Platform admins. None of this is unique to Microsoft. Gartner predicts that [over 40% of agentic AI projects will be canceled by the end of 2027](https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027), citing escalating costs, unclear business value, and inadequate risk controls, and it warns about "agent washing," vendors rebranding old chatbots as agents. Copilot Studio is a real platform, not agent washing, but the same discipline applies: start with one clear, valuable use case, measure it, and expand only when it pays off. ## What Copilot Studio Means for Your Brand's AI Visibility There is one more angle worth covering, and it matters whether or not you ever build an agent yourself. Every Copilot Studio agent runs on content. The internal ones you build are only as good as the knowledge you feed them, which is a direct argument for keeping your documentation clean, current, and well-structured. Vague source content produces vague agents. The more interesting side is the agents you do not control. When a partner, a customer, or Microsoft 365 Copilot itself grounds an answer on the open web, your public content is the raw material. Copilot's consumer surfaces lean on the Bing index, so the same work that earns you a citation there strongly influences whether an agent names your brand or a competitor's. That is the heart of [getting cited in Microsoft Copilot](https://geotoolbox.ai/blog/copilot-seo), and it is why an [agent-ready website](https://geotoolbox.ai/blog/agent-ready-website) is becoming a visibility surface in its own right. In our experience, the brands that show up well in AI answers are not the ones with the cleverest agents. They are the ones whose content is reachable and structured cleanly enough for any agent to read and lift. That is the boring, durable lever underneath all the agent hype. If you want to know whether agents and AI crawlers can actually reach and read your content, that is exactly what we built our tools to check. Run your site through the [AI Readiness](https://geotoolbox.ai/tools/ai-readiness) scan to check whether AI agents can reach, crawl, and parse your pages, and use the [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) to confirm you are not blocking the major AI crawlers before they ever get to your content. ## Frequently Asked Questions ### Is Copilot Studio free? Only for evaluation and for licensed internal use. The trial builds and tests agents but cannot publish them, and internal use by Microsoft 365 Copilot license holders is included at no extra credit cost. Any external, autonomous, or unlicensed use spends Copilot Credits. ### What is the difference between Copilot Studio and Microsoft Copilot? Microsoft 365 Copilot is the finished assistant employees use inside Word, Excel, and Teams. Copilot Studio is the low-code platform you use to build your own custom agents. ### Is Copilot Studio the same as Power Virtual Agents? Yes. Power Virtual Agents was renamed Copilot Studio in late 2023 and expanded from chatbots into full AI agents. If you built a PVA bot, it now lives in Copilot Studio. ### Is Copilot shutting down? No. The confusion comes from Power Virtual Agents being renamed and folded into Copilot Studio, not removed. The only sunset is the legacy Copilot Studio for Teams app, which stopped creating classic chatbots on June 30, 2026 and now redirects makers to the web app. ### How much does a Copilot Studio agent cost? It depends entirely on what the agent does. A scripted answer costs 1 Copilot Credit, a generative answer costs 2, and a reasoning response is metered per token and costs many times more. A prepaid pack is $200 a month for 25,000 credits, so the same agent can run a few dollars or several hundred depending on its design. ### Do you need a license to use Copilot Studio, and who needs one? The maker license is free, but your tenant needs a Copilot Credit pack subscription in place first. End users do not need their own Copilot Studio license; only the people building agents do, and a Microsoft 365 Copilot license covers internal use. ### Can you build a Copilot Studio agent without coding? For simple agents, yes, you describe what you want in plain language. More advanced agents tend to require variables, expressions, JSON, and connector setup, so "low-code" is accurate but "no-code" overstates it. ## Sources - Copilot Studio overview, Microsoft Learn - `learn.microsoft.com/en-us/microsoft-copilot-studio/fundamentals-what-is-copilot-studio` - Copilot Studio billing rates and management, Microsoft Learn - `learn.microsoft.com/en-us/microsoft-copilot-studio/requirements-messages-management` - Copilot Studio licensing, Microsoft Learn - `learn.microsoft.com/en-us/microsoft-copilot-studio/billing-licensing` - Microsoft Copilot Studio pricing - `microsoft.com/en-us/microsoft-365-copilot/pricing/copilot-studio` - Copilot Studio pay-as-you-go pricing, Microsoft Azure - `azure.microsoft.com/en-us/pricing/details/copilot-studio` - Copilot Studio agent usage estimator, Microsoft - `microsoft.github.io/copilot-studio-estimator` - National Zakat Foundation customer story, Microsoft - `microsoft.com/en/customers/story/23068-national-zakat-foundation-microsoft-copilot-studio` - Gartner: over 40% of agentic AI projects will be canceled by the end of 2027 - `gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027` - r/copilotstudio community discussion, Reddit - `reddit.com/r/copilotstudio/comments/1ip7y97/what_is_your_experience_with_copilot_studio_does` --- ## Microsoft Copilot vs Google Gemini: Which Is Better? (2026) > An honest, current comparison of Microsoft Copilot and Google Gemini: models, pricing, coding, privacy, and which AI assistant actually fits your work in 2026. - Canonical: https://geotoolbox.ai/blog/copilot-vs-gemini - Published: 2026-06-28 · Updated: 2026-08-08 Copilot vs Gemini usually comes down to one question that has nothing to do with which AI is "smarter": which software do you already use all day? Microsoft Copilot is built into Windows and Office; Google Gemini is built into Android and Google Workspace. Both write, code, research, and generate images, and after a year of fast updates the feature gap is small. This is an honest, current comparison of the two as of July 2026, covering the models behind each, what they cost, who wins for coding, writing, research, and meetings, how they handle your data, and the one thing neither comparison usually mentions: how to get your own brand recommended when people ask these assistants for advice. ## Copilot vs Gemini at a Glance The short version: Microsoft Copilot wins if your work lives in Windows and Microsoft 365, and Google Gemini wins if it lives in Google Workspace or you need heavy research, long documents, and video. Almost everything else follows from that one fact. The differences that matter now are which suite they plug into, which underlying models they run, how they price, and how they treat your data.
DimensionMicrosoft CopilotGoogle Gemini
MakerMicrosoftGoogle
Lives inWindows, Microsoft 365 (Word, Excel, Outlook, Teams)Android, Google Workspace (Gmail, Docs, Sheets, Meet)
Default chat model (July 2026)OpenAI's GPT-5.6 family, with Anthropic Claude as an optionGemini 3.6 Flash, with 3.1 Pro and Deep Think for harder tasks
Strongest atOffice documents, meetings, enterprise data groundingResearch, large files, image and video generation
Free tierYes (Copilot Chat)Yes (Gemini app)
Main paid consumer planMicrosoft 365 Premium, $19.99/moGoogle AI Pro, $19.99/mo (AI Plus from $4.99)
Business AIMicrosoft 365 Copilot, around $30/user/moGemini in Workspace, about $14-22/user/mo
![Scorecard comparing Microsoft Copilot and Google Gemini in July 2026: Copilot wins Office documents, meetings, and grounding answers in your company's data; Gemini wins research and long documents, image and video, and the cheapest paid entry; writing and coding are roughly tied.](/blog/copilot-vs-gemini/copilot-vs-gemini-scorecard.png)
Each assistant wins the jobs that live in its own ecosystem; writing and coding are close enough to call a tie.
If you already know which ecosystem you live in, jump to the pricing and privacy sections below. The model rosters move almost monthly, so check the update date at the top of this page before you act on anything here. ## What You're Actually Comparing Before the head-to-head, clear up the names. "Copilot" and "Gemini" each refer to a family of products, and people argue past each other because they are comparing different members of those families. On the Microsoft side there are three things called Copilot: - **Microsoft Copilot** is the free consumer chatbot in Windows, Edge, and the Copilot app. It searches the web and reads files you upload, but it does not see your company's data. - **Microsoft 365 Copilot** is the paid business add-on. This is the one that reads your emails, meetings, and files through Microsoft Graph and works inside Word, Excel, and Teams. When Microsoft markets "Copilot for work," this is what it means. We cover what that is in our [guide to Microsoft Copilot](https://geotoolbox.ai/blog/what-is-copilot). - **GitHub Copilot** is a separate developer product that lives in your code editor. It shares a name and an owner with the others, but it is bought, priced, and used on its own. On the Google side the split is similar: - **The Gemini app** is the free consumer assistant on the web and on Android and iOS. - **Gemini for Workspace** brings Gemini into Gmail, Docs, Sheets, and Meet for individuals and businesses. We break down the lineup in our [guide to Google Gemini](https://geotoolbox.ai/blog/what-is-gemini). - **Gemini Code Assist** (plus the Gemini CLI and Google's Antigravity editor) is the developer-facing side, the direct rival to GitHub Copilot. Most people typing "copilot vs gemini" mean the everyday assistants and their work versions, so that is what this comparison centers on. We give coding its own section, because the developer tools answer a different question. ## The Models and Tech Behind Each The rosters move fast, so here is where they stand in July 2026. Copilot is not a single model. It is an orchestration layer that routes your request to the best engine. By default it runs on OpenAI's latest GPT-5.6 models, and where an admin has enabled it, [Microsoft 365 Copilot users can select Anthropic's Claude](https://www.microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot/) from the model picker in Copilot Chat, the Researcher agent, and Copilot Studio. Microsoft has also begun adding its own in-house MAI models to the mix. The newer agent layer, [Copilot Cowork](https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/), which runs multi-step jobs in the background, became generally available on June 16, 2026. It launched on Anthropic's Claude Opus 4.8 and Sonnet 4.6, with Sonnet 5 replacing Sonnet 4.6 on July 2, 2026, and OpenAI's GPT-5.6 became its preferred model on July 9, 2026. Gemini runs on Google's own models end to end. As of July 2026 the default is **Gemini 3.6 Flash** for speed, with **Gemini 3.1 Pro** and **3.1 Deep Think** handling heavier reasoning, per [Google DeepMind](https://deepmind.google/models/gemini/); a higher-end [Gemini 3.5 Pro](https://geotoolbox.ai/blog/gemini-3-5-pro) is coming soon. Around it sit Google's creative engines: Veo for video, Lyria for music, and Nano Banana for images. At its 2026 developer conference Google also previewed [Gemini Spark](https://blog.google/innovation-and-ai/technology/ai/google-io-2026-all-our-announcements/), an agent that acts across your apps, and a more multimodal Omni layer. The one technical gap that still matters is context. Gemini handles up to roughly a million tokens of explicit context, so you can paste a long contract, a full research report, or a stack of documents into one prompt. The consumer Copilot chat sits closer to 128K tokens. Copilot makes up for that a different way: in the business version, Microsoft Graph quietly pulls in the relevant email, file, or meeting without you uploading anything. One approach wins on raw volume you control; the other wins on ambient awareness of your work.
CapabilityMicrosoft CopilotGoogle Gemini
Default chat modelGPT-5.6 family (routed)Gemini 3.6 Flash
Heavy reasoningGPT-5.6 (deep-reasoning mode); Claude optionGemini 3.1 Pro, 3.1 Deep Think
Other modelsAnthropic Claude (Opus / Sonnet); Microsoft MAIVeo (video), Lyria (music), Nano Banana (image)
Agent layerCopilot Cowork (Claude Opus 4.8 / Sonnet 5; GPT-5.6 now preferred)Gemini Agent Mode, Gemini Spark
Explicit context~128K tokens (consumer chat)~1 million tokens
Ambient contextMicrosoft Graph (mail, files, meetings)Large file and multi-file uploads
Native video generationNo (images via OpenAI's GPT-4o)Yes (Veo)
## Head-to-Head by Job No single tool wins every task. Here is how they split, based on independent hands-on reviews and our own testing. ### Writing and Content Gemini is the stronger creative writer. It tends to produce more natural, varied prose, which is why marketers and content teams lean on it for drafts, social copy, and brainstorming. Copilot is the stronger business writer: tighter on structured documents, executive emails, and reports, and it can turn a Word doc into a formatted slide deck with speaker notes in one step. The twist is that Copilot lets paid users swap in Anthropic's Claude, widely regarded as one of the most natural writers among frontier models. So "who writes better" partly depends on which model you point Copilot at. For everyday drafting, call it a tie that tilts to Gemini for creative work and Copilot for polished business output. ### Research Gemini has the clearer edge for open-web research and long-document analysis. Its Deep Research mode autonomously reads dozens of live pages and returns a structured brief, the million-token context lets you load entire documents in one pass, and NotebookLM turns your own sources into a study-ready notebook. In [ZDNET's hands-on test](https://www.zdnet.com/home-and-office/work-life/gemini-vs-copilot/), Gemini produced an accurate multi-city train itinerary while Copilot got it wrong, citing direct connections that did not exist and a map that misplaced whole cities. Copilot's research edge is narrower but real: inside a business, it grounds answers in your own files and meetings through Microsoft Graph, which Gemini cannot do for Microsoft data. For open-web research, reach for Gemini. For "what did we decide in last week's planning meeting," Copilot is the only one that can answer. ### Coding This is the one place the everyday-assistant comparison breaks down, because the real contest is between **GitHub Copilot** and **Gemini Code Assist**, and they suit different developers. GitHub Copilot remains the smoother day-to-day tool: fast inline completions, mature Visual Studio and VS Code integration, and a tight GitHub workflow. Gemini Code Assist (with the Gemini CLI and Google's Antigravity editor) leans on that million-token context to reason across an entire repository, explain unfamiliar code, and debug multi-file problems. Developers tend to report the same split: Gemini can hold a whole project in context, while GitHub Copilot works best when you feed it a file at a time. Pick GitHub Copilot for shipping features quickly in an existing workflow; pick Gemini for understanding a large or unfamiliar codebase. ### Office vs Workspace This is decided by where your documents already live, and the reality is that each tool dominates its own suite. Copilot edits natively inside Word and Excel, building formulas and reformatting tables in the document you are already in. Gemini does the same inside Docs, Sheets, and Gmail. Microsoft's own [comparison page](https://www.microsoft.com/en-us/microsoft-365-copilot/copilot-vs-gemini-enterprise) leans hard on this, claiming Gemini "creates a new file instead of updating" when asked to work on an Office document. That is true of Gemini reaching into Microsoft's files, and it is Microsoft's framing of its own test, not a neutral finding. Inside Google's own apps, Gemini edits Docs and Sheets just as natively. The takeaway is not that one tool is broken; it is that neither reaches cleanly into the other company's file formats. Stay in your suite and both work well. ### Image and Video Gemini is ahead on creative media. It generates images through Nano Banana and is one of the few assistants with native video generation through Veo, useful for anyone producing visual content. Copilot handles images with OpenAI's GPT-4o image generation and added Copilot 3D for turning images into simple 3D models, but it has no native video generation. If visuals are central to your work, Gemini is the more complete creative studio. ### Meetings Copilot wins meetings, mostly because of Teams. It transcribes and summarizes calls in real time, tracks action items and owners, and can recap up to a month of chat history. Gemini summarizes inside Google Meet but the experience is lighter. If your team runs on Teams, this alone can decide the comparison. ## Pricing Compared (2026) Headline consumer pricing is nearly identical: both charge $19.99 a month for their main paid plan. The differences are at the edges, and the labels changed in late 2025, which is where the confusion starts. On the Microsoft side, the standalone Copilot Pro plan is gone. Microsoft [retired it on October 1, 2025 and folded its AI features into Microsoft 365 Premium](https://www.microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse/) at the same $19.99, which now also bundles the Office apps and storage. Existing Copilot Pro subscribers kept access until [support ended on August 1, 2026](https://support.microsoft.com/en-us/microsoft-365-copilot/about-microsoft-copilot-pro). For deeper detail see our [Microsoft Copilot pricing guide](https://geotoolbox.ai/blog/copilot-pricing). On the Google side, [Google's plan lineup](https://one.google.com/about/google-ai-plans/) added a cheaper entry in 2026: Google AI Plus at $4.99 a month sits below AI Pro at $19.99, and AI Ultra starts at $99.99 for heavy users. Our [Gemini pricing guide](https://geotoolbox.ai/blog/gemini-pricing) breaks down what each tier includes.
Plan levelMicrosoft CopilotGoogle Gemini
FreeCopilot (Chat) - $0Gemini app - $0
Entry paidMicrosoft 365 Personal - $9.99/moGoogle AI Plus - $4.99/mo
Main consumerMicrosoft 365 Premium - $19.99/moGoogle AI Pro - $19.99/mo
Top consumer(folded into Premium)Google AI Ultra - from $99.99/mo
BusinessMicrosoft 365 Copilot - $30/user/mo (or $21 Business, $18 promo through Sep 30 2026)Workspace plans that bundle Gemini - about $14-22/user/mo; Enterprise custom
Two things to watch on the business side. Microsoft is raising its base Microsoft 365 plan prices on July 1, 2026, though [the Microsoft 365 Copilot add-on price](https://www.microsoft.com/en-us/microsoft-365-copilot/pricing) holds steady, so a Copilot rollout may cost more than last year through the underlying license. And Google now bundles Gemini into several Workspace tiers, so many businesses already have it without a separate add-on. Before you buy either, check what your existing subscription already includes; a lot of teams are paying for AI they already own. For personal use, the free tiers cover most everyday tasks, so pay only when you hit usage limits or need the top models, longer context, or video. ## Privacy and Data Training Privacy is the one area where the two are genuinely different, and it depends entirely on whether you are a consumer or a business. For **business and enterprise** tiers, the two are close. Microsoft 365 Copilot keeps prompts and company data inside your tenant under Enterprise Data Protection and [does not use them to train its foundation models](https://learn.microsoft.com/en-us/microsoft-365/copilot/microsoft-365-copilot-privacy). Gemini for Workspace and Gemini Enterprise make the same promise: your content stays governed by Workspace controls and is not used for training. On governance, Copilot leans on Microsoft Purview sensitivity labels and data-loss-prevention policies, which Microsoft positions as an enterprise edge; Google offers its own Workspace admin and DLP controls. If you are buying for a company, the privacy posture is much closer than the headlines suggest. For **consumers**, the gap is real, and it is where the "Gemini red flag" question comes from. On free and paid individual Gemini plans, Google treats your chats as consumer data: a subset is [reviewed by humans and used to improve its models](https://support.google.com/gemini/answer/13594961), and the main control is turning off Gemini Apps Activity, which also stops new chats from being saved to your history. Paying for AI Pro buys you better models, not better privacy. The concern sharpened in late 2025 and 2026 over Gemini's reach into Gmail. A [lawsuit, Thele v. Google](https://natlawreview.com/article/silent-switch-new-lawsuit-alleges-google-uses-gemini-ai-secretly-read-gmail-chat), alleges Google used an existing "Smart Features" toggle to switch on deeper Gemini data access without a clear, separate consent prompt; as of mid-2026 the case is at an early stage, and Google disputes the claims, saying it did not change any setting to train on Gmail. Treat it as an open allegation, not a settled fact, but it is a fair reason to check your settings. Microsoft's consumer Copilot is not automatically cleaner. Free Copilot interactions can be processed under standard Bing terms unless you are signed in with a work or school account. The summary: for either tool, business tiers protect your data and consumer tiers ask you to manage your own settings. If privacy is the deciding factor and you are an individual, read the controls before you commit. ## Which Should You Choose? The decision is less about which model is smarter this month and more about where you already work. Match your situation to the table.
If you...Choose
Live in Word, Excel, Outlook, and TeamsMicrosoft Copilot
Live in Gmail, Docs, Sheets, and MeetGoogle Gemini
Do heavy research or work with very long documentsGoogle Gemini
Need meeting transcription and follow-upMicrosoft Copilot (Teams)
Produce images or videoGoogle Gemini
Want your company's own data in answersMicrosoft 365 Copilot
Ship code daily in an existing GitHub workflowGitHub Copilot
Need to reason across a large codebaseGemini Code Assist
Want the cheapest paid entry pointGoogle AI Plus ($4.99)
For many people, the answer is "use both." The free tiers cost nothing, so a common split is Gemini for research, long documents, and creative work and Copilot for Office files and meetings, rather than forcing every task through one assistant. If ChatGPT is on your shortlist too, it is the neutral third option for anyone not committed to either suite; we compare it in [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt) and [Microsoft Copilot vs ChatGPT](https://geotoolbox.ai/blog/microsoft-copilot-vs-chatgpt). In our experience helping brands track AI assistants, the "which is better" question matters far less than people expect, because most users never switch ecosystems for an AI tool. They keep the suite they already pay for and use whichever assistant ships with it. Because of that inertia, the more useful question for a business is not which tool to use, but whether these assistants can find and recommend you when your customers ask them. ## How to Get Cited by Copilot and Gemini When someone asks Copilot or Gemini "what's the best tool for X," both name specific brands and link to sources. If your competitors show up in that answer and you do not, you are losing customers before they ever reach a search results page. The mechanics of getting named are different for each, because they read different parts of the web. Copilot grounds its web answers on the **Bing index**. Microsoft's [Bing Webmaster Guidelines](https://www.bing.com/webmasters/help/webmaster-guidelines-30fba23a) state that Bing and Copilot search rely on the same crawling, indexing, and ranking foundation. So the entry ticket to a Copilot citation is being well-indexed by Bing, not just Google. Our [Copilot SEO guide](https://geotoolbox.ai/blog/copilot-seo) covers the specifics. Gemini grounds on the **Google index** and the same content feeds [Google's AI features](https://developers.google.com/search/docs/appearance/ai-features), including [AI Overviews](https://geotoolbox.ai/glossary/ai-overviews) and [AI Mode](https://geotoolbox.ai/glossary/google-ai-mode). Being cited by Gemini and by Google's AI surfaces is largely the same job: earn a place in Google's index and write passages an AI can lift cleanly. Our [Gemini SEO guide](https://geotoolbox.ai/blog/gemini-seo) goes deeper. The work that pays off in both engines overlaps. Write answer-first, self-contained passages that state a fact and back it up, because that is what these assistants quote. Earn corroboration from the third-party sources they trust, since both engines cross-check claims against the wider web rather than taking your word for it. And keep your entity and facts consistent across the sites that describe you. This is the same discipline behind [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) generally, and it matters more as [agentic assistants](https://geotoolbox.ai/blog/agentic-ai) start acting on these answers, not just displaying them. This is the gap most Copilot-versus-Gemini comparisons miss entirely. They tell you which assistant to use; they never tell you how to be the brand those assistants recommend. That second question is where the commercial value is, and it is what our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) was built to answer: it queries the engines your customers actually use, including Gemini and Bing Copilot, shows which sources each one cites in your category, and flags where your brand is missing from those answers. ## Frequently Asked Questions ### Is Gemini better than Copilot? Neither is better across the board. Gemini wins on research, long documents, and image and video. Copilot wins on Microsoft 365 work, meetings, and grounding answers in your company's own data. The right pick is the one that lives in the suite you already use. ### What can Microsoft Copilot do that Google Gemini can't? Three things stand out. It edits natively inside Word and Excel, it can read your organization's emails, files, and meetings through Microsoft Graph (in the paid business version), and it lets you switch the underlying model to Anthropic's Claude. Gemini cannot reach into Microsoft 365 data the way Copilot does. ### Why is Copilot shutting down? It is not. The confusion comes from Microsoft retiring the standalone Copilot Pro consumer plan on October 1, 2025 and folding its features into Microsoft 365 Premium. Existing Copilot Pro subscriptions kept working until support ended on August 1, 2026, and the business Microsoft 365 Copilot your employer deploys is a separate product that is unaffected. ### Is Microsoft Copilot just ChatGPT? No. Copilot uses OpenAI's GPT-5.6 family by default, the same model family behind ChatGPT, but it adds Microsoft Graph grounding, Anthropic Claude as a switchable option, Microsoft's own MAI models, and deep Office integration. It is an orchestration layer over several models, not a rebadged ChatGPT. ### Why does Copilot feel weaker than ChatGPT if they use the same models? They share the GPT-5.6 family, but the experience differs. Consumer Copilot wraps the model in heavier web grounding and tighter guardrails and can route simpler prompts to lighter models, so on open creative or reasoning tasks raw ChatGPT can feel more capable. Microsoft 365 Copilot trades some of that raw feel for something ChatGPT cannot do: answer using your own company's files and meetings. ### Why do people call Gemini a privacy red flag? Because on free and paid consumer plans, Google treats your chats as consumer data that can be reviewed by humans and used to improve its models unless you turn off Gemini Apps Activity. A 2026 lawsuit also alleges Google enabled deeper Gmail access without clear consent, a claim Google disputes and that is still in court. Business and enterprise Gemini tiers do not train on your data. ### Can I use Copilot and Gemini together? Yes, and many people do. The free tiers cost nothing, so a common setup is Gemini for research and creative drafting and Copilot for Office documents and meetings. You do not have to commit to one ecosystem to benefit from both. ## The Bottom Line Microsoft Copilot and Google Gemini are both strong, and after a year of rapid updates they do most of the same jobs. The choice comes down to where your work already lives, with Gemini pulling ahead on research, large files, and creative media, and Copilot pulling ahead on Office work, meetings, and grounding answers in your own data. Pick the one that matches your suite, keep the other open for the tasks it does better, and check the privacy settings either way. For brands, the more important shift is that buyers now ask these assistants for recommendations directly. Whether Copilot and Gemini name you, link you, or skip you is becoming its own visibility channel. If you want to see which sources they cite in your category and where your brand is absent, that is exactly what [geotoolbox's Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) is for. ## Sources - Microsoft 365 blog: Expanding model choice in Microsoft 365 Copilot - `microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot` - Microsoft 365 blog: Meet Microsoft 365 Premium - `microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse` - Microsoft 365 blog: Copilot Cowork is now generally available - `microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available` - Microsoft Support: About Microsoft Copilot Pro - `support.microsoft.com/en-us/microsoft-365-copilot/about-microsoft-copilot-pro` - Microsoft 365 Copilot pricing - `microsoft.com/en-us/microsoft-365-copilot/pricing` - Microsoft Learn: Data, privacy, and security for Microsoft 365 Copilot - `learn.microsoft.com/en-us/microsoft-365/copilot/microsoft-365-copilot-privacy` - Bing Webmaster Guidelines - `bing.com/webmasters/help/webmaster-guidelines-30fba23a` - Microsoft: Copilot vs Gemini Enterprise (Microsoft's own comparison) - `microsoft.com/en-us/microsoft-365-copilot/copilot-vs-gemini-enterprise` - Google DeepMind: Gemini models - `deepmind.google/models/gemini` - Google: Google AI plans and pricing - `one.google.com/about/google-ai-plans` - Google blog: Google I/O 2026 announcements - `blog.google/innovation-and-ai/technology/ai/google-io-2026-all-our-announcements` - Google Search Central: AI features and your website - `developers.google.com/search/docs/appearance/ai-features` - Google Help: Gemini Apps privacy - `support.google.com/gemini/answer/13594961` - ZDNET: Gemini vs Copilot, 7 everyday tasks - `zdnet.com/home-and-office/work-life/gemini-vs-copilot` - PCWorld: Microsoft adds 365 Premium, cuts Copilot Pro - `pcworld.com/article/2925962/microsoft-adds-microsoft-365-premium-cuts-copilot-pro.html` - National Law Review: lawsuit alleges Gemini read Gmail, Chat, and Meet (Thele v. Google) - `natlawreview.com/article/silent-switch-new-lawsuit-alleges-google-uses-gemini-ai-secretly-read-gmail-chat` --- ## Gemini 3.5 Pro: Release Date, Specs & What's Confirmed (2026) > Gemini 3.5 Pro is not out yet. A dated look at what Google confirmed, what's only rumored (2M context, Deep Think, pricing), and what it means for AI search. - Canonical: https://geotoolbox.ai/blog/gemini-3-5-pro - Published: 2026-06-28 · Updated: 2026-08-08 Gemini 3.5 Pro is the Google model everyone is waiting for and almost no one can use. As of August 8, 2026 it still has not shipped, Google has published no specs for it, and most of what you will read about its context window, pricing, reasoning, and release date is rumor formatted to look official. This is a dated tracker that separates the two. One disambiguation first: Gemini 3.5 Pro is the unreleased Pro tier, separate from Gemini 3.5 Flash, which is live, and Gemini 3.1 Pro, which it succeeds in the lineup. ## Is Gemini 3.5 Pro Out Yet? No. As of August 8, 2026, Gemini 3.5 Pro is still not generally available, and most of what you can read about its specs is guesswork. (Last verified August 8, 2026. When the rollout completes, the API model-list status line covered below is the fastest way to check.) Google announced the Gemini 3.5 family at [Google I/O on May 19, 2026](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/), but only one model in that family actually shipped: Gemini 3.5 Flash. About Pro, Google said one substantive thing, "It's already being used internally, and we look forward to rolling it out next month." No model card, no specs, no pricing, no firm date. The video side of that same keynote, [Gemini Omni](https://geotoolbox.ai/blog/gemini-omni), did ship. Three names get tangled here, so worth separating them up front: - **Gemini 3.5 Flash** is live, though on July 21, 2026 its successor [Gemini 3.6 Flash](https://geotoolbox.ai/blog/gemini-3-6-flash-vs-3-5-flash-lite) took over as the default model in the Gemini app and AI Mode in Search. - **Gemini 3.5 Pro** is the one everyone is waiting for. It has not shipped. - **Gemini 3.1 Pro** is the current top-tier Pro model, listed as preview. It is the real predecessor to 3.5 Pro, not Gemini 2.5 Pro, which a lot of coverage gets wrong. You can confirm the status yourself. The [Gemini API model list](https://ai.google.dev/gemini-api/docs/models) shows `gemini-3.5-flash` marked **Stable**, and there is no `gemini-3.5-pro` entry at all. The next tier up is still Gemini 3.1 Pro. When a model exists on Google's API, it has an ID and a launch stage. Pro has neither yet. On timing, the story got worse, not better. On July 16, 2026, Alphabet shares [fell about 4%](https://www.cnbc.com/2026/07/16/alphabet-stock-gemini-3-5-pro-ai.html) after Bloomberg reported that Pro is running months behind schedule. The reported cause is more than a testing pause: Pro's coding performance came in short of Google's own internal expectations, updated training data reportedly failed to close the gap, and rivals like OpenAI and Anthropic have pulled ahead on code. A Google spokesperson said only that the company is "currently testing 3.5 Pro," neither denying the delay nor giving a new date. Earlier reporting had pointed to a July 17, 2026 target and a full architectural rebuild of the scrapped 2.5 Pro base model, aimed at stronger math reasoning, SVG generation, and image quality to keep pace with GPT-5.6 and Fable 5. That July 17 date has now passed with no release, and there is still no model card, no API entry, and no pricing. Treat every date here the same way, plausible, widely repeated, and unverified until a model card exists. For the full picture of where Pro sits in [Google's Gemini lineup](https://geotoolbox.ai/blog/what-is-gemini), the short version is: announced, delayed, not arrived.
![Gemini 3.5 rollout status: Flash generally available, Pro in internal preview, GA reported months behind schedule.](/blog/gemini-3-5-pro/gemini-3-5-rollout-status.png)
Where each Gemini 3.5 model stands: Flash shipped, Pro delayed and still internal-only.
## What Google Has Actually Confirmed Here is the base to reason from. Almost everything Google has stated about Gemini 3.5 is about Flash, and Flash is the same generation as Pro, so its confirmed numbers are the most reliable signal we have for what Pro will be built on. From [Google's own Gemini 3.5 announcement](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5/), these are not rumors: - **Gemini 3.5 Flash is generally available** and was the default model in the Gemini app and AI Mode in Search globally at the time of that announcement. (Google later shipped Gemini 3.6 Flash on July 21, 2026, which took over as the default.) - Flash posts real benchmark scores: **76.2% on Terminal-Bench 2.1**, **1656 Elo on GDPval-AA**, **83.6% on MCP Atlas**, and **84.2% on CharXiv Reasoning** for multimodal understanding. - Google states Flash is **outperforming Gemini 3.1 Pro on its coding and agentic benchmarks**, which is the headline: a Flash-tier model beating the previous Pro tier on Google's chosen hard tasks. - Gemini 3.5 Pro is **being used internally** and was set to roll out the following month. That is the confirmed set that matters for sizing up Pro. Notice what is missing: there is no Gemini 3.5 Pro model card, so its context window, its reasoning modes, its pricing, and its benchmark scores are all officially unstated. Pro relates to Flash the way it always has: Flash is the fast, cheaper tier, Pro the slower, more expensive, higher-ceiling one. Given Flash's results against the old Pro tier, a 3.5 Pro that clears Flash would be a genuine step up. But "would be" is the operative phrase. Until Google ships a model card, every Pro specification is an estimate, and the next section is where the estimates start getting presented as facts. ## Gemini 3.5 Pro: Rumored vs Confirmed This is where most coverage quietly fails. Specs that Google has never published get listed in clean "specifications" tables as if they were official. Here is the same data with the source attached.
SpecWhat's claimedWho's claiming itGoogle official?
Context window2 million tokensThird-party explainersNo. Google has published no Pro token count; the only 1M figure in circulation is for Flash, not Pro
Deep Think reasoningA dedicated extended-reasoning modeThird-party blogsNo. Not in any Google model card for 3.5 Pro
API pricing~$15 input / $60 output per million tokens (roughly 7 to 10x Flash)ByteIota and others, labeled "estimated"No. No published rate card. Others assume it lands near current Gemini 3.1 Pro pricing, a wide spread
Deep Think accessGated to a ~$250/mo Google AI Ultra tierThird-party reportingNo. Unconfirmed
ReleaseDelayed several months; the July 17, 2026 target passed with no releaseBloomberg, via CNBC (July 16, 2026)No. Google confirms only that it is "currently testing" Pro; no new date, model card, or API entry
ArchitecturePrior 2.5 Pro base scrapped for a full rebuild and fresh pre-training (math, SVG, image quality)Third-party reportingNo. Not confirmed by Google
The cleanest example is the [context window](https://geotoolbox.ai/glossary/context-window). Several widely-shared specs tables list "2 million tokens" for Gemini 3.5 Pro as a confirmed figure. Google's actual announcement gives no token count for Pro at all, and the 1M figure in circulation is associated with Flash, not Pro. So a spec presented as fact is, at the source level, an assumption. Pricing tells the same story from a different angle. [ByteIota's estimate](https://byteiota.com/gemini-3-5-pro-2m-tokens-deep-think-and-the-10x-pricing-problem/) puts Pro at roughly $15 input and $60 output per million tokens, about seven to ten times Flash, and is careful to label it an estimate. Others simply assume Pro will land near [current Gemini 3.1 Pro pricing](https://geotoolbox.ai/blog/gemini-api-pricing), several times lower. When credible guesses disagree by 5x or more, that is your signal that no one has the real number. The [Deep Think reasoning mode](https://geotoolbox.ai/glossary/reasoning-model) is real as a concept Google has discussed elsewhere, but it has not been attached to a published Gemini 3.5 Pro spec. Treat it as possible but unconfirmed until Google ties it to a 3.5 Pro model card. ## Gemini 3.5 Pro vs 3.5 Flash vs 3.1 Pro Get the lineage right and the comparison gets simpler. The model Gemini 3.5 Pro replaces is **Gemini 3.1 Pro**, not Gemini 2.5 Pro. That older "vs 2.5 Pro" framing shows up everywhere and skips a whole generation.
ModelStatusConfirmed benchmarksWhere it fits
Gemini 3.5 FlashGenerally available; was the app + AI Mode default until 3.6 Flash (July 2026)Terminal-Bench 2.1 76.2%, GDPval-AA 1656, MCP Atlas 83.6%, CharXiv 84.2%Fast, cheaper tier you can use today; already beats 3.1 Pro on coding/agentic
Gemini 3.5 ProNot released; internal preview, reported months behind schedule (Bloomberg, July 2026)None publishedThe higher-ceiling tier, on paper. No verified numbers exist yet
Gemini 3.1 ProPreview (current top Pro tier)Per Google, now bettered by 3.5 Flash on coding/agenticThe real predecessor and today's Pro option until 3.5 Pro ships
The read: the only 3.5 model you can actually compare on data is Flash, and it looks strong. Pro's row is empty by necessity. Anyone publishing a Gemini 3.5 Pro benchmark table right now is extrapolating from Flash or inventing numbers, because Google has released none. You will also see community estimates of Pro's parameter count framed as if size settles the question. It does not. Flash already beating the previous Pro tier on hard tasks is the clearest evidence that architecture and training matter more than raw size this generation. If you are choosing a model today, the practical comparison is Flash against the alternatives you already use, which is the ground covered in [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt) and [Claude vs Gemini](https://geotoolbox.ai/blog/claude-vs-gemini), not a Pro model you cannot run. ## How to Access Gemini 3.5 Pro (and Why You Might Not Be Able to Yet) You cannot, not directly, and the reason is structural. A model needs a public ID and a launch stage before you can call it on the public Gemini API, and Gemini 3.5 Pro has neither. The [Gemini API model list](https://ai.google.dev/gemini-api/docs/models) carries `gemini-3.5-flash` as Stable and simply has no `gemini-3.5-pro` line. What you can use right now is Flash, and Google has put it across its main AI surfaces: - The **Gemini app** and **AI Mode in Search**, where 3.6 Flash became the default in July 2026, succeeding 3.5 Flash - **Google AI Studio** and the **Gemini API**, calling `gemini-3.5-flash` When Pro does arrive, expect the usual path rather than a flip-the-switch global launch. If it follows the pattern of past Gemini previews, it surfaces first in AI Studio and Vertex AI as an allowlisted preview, often United States first, before reaching the consumer app and broad API access. That is why "it's rolling out" and "I can use it" are weeks to months apart, and why teams with data-residency requirements outside the US should plan for a later usable date than the headline announcement implies. In practice that means watching for `gemini-3.5-pro` to appear in the Model Garden on Vertex AI, now branded the Gemini Enterprise Agent Platform, and the AI Studio model picker, and not hardcoding the ID until it lists a stable launch stage. If your workflow depends on a specific model being callable in production, track the launch stage on the API model page, not the blog announcement. "Generally available" on the model list is the line that matters. A model that is "announced" or even "in preview" can still be pulled, throttled, or region-locked. ## What Gemini 3.5 Pro Means for AI Search Visibility Here is the part that affects your traffic whether or not Pro ever ships on your timeline: the Gemini Flash shift is already live in Google's AI surfaces. Google says the current Flash model, now Gemini 3.6 Flash (which succeeded 3.5 Flash as the default on July 21, 2026), is the default in [AI Mode](https://geotoolbox.ai/glossary/google-ai-mode), and AI Overviews, the AI answers in regular Search, run on Gemini models too. The model upgrade reached your audience the day it shipped, not on Pro's launch day. That surface is not niche. At I/O, [Sundar Pichai said AI Overviews has over 2.5 billion monthly users](https://www.eweek.com/news/google-io-gemini-agentic-ai-era-2026/). And those answers change behavior. Pew Research found that [when an AI summary appears, people click a traditional result in only 8% of visits, versus 15% without one](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/), and they click a link inside the summary in just 1% of visits. More of the click is staying inside the answer rather than reaching your page, and those answers are generated by Gemini models. So the question is not "should I migrate to Gemini 3.5 Pro." It is "when Google's AI summarizes my topic, does it pull from my page." That is a retrieval problem, and it rewards specific habits: state facts in clean, liftable sentences, put the real answer near the top of each section, use question-shaped headings, and keep an FAQ that answers the obvious follow-ups directly. In our experience scanning sites for AI visibility, which is what geotoolbox does, the pages that get cited are rarely the longest ones. They tend to be the ones that state a fact plainly enough that a model can lift a single sentence without rewriting it. That is the same discipline whether the engine is Flash today or Pro next quarter, and it is the core of [getting cited in AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) and the wider practice of [Gemini SEO](https://geotoolbox.ai/blog/gemini-seo). ## What People Are Hitting on 3.5 Flash While They Wait Since Pro does not exist yet, most people reaching for a current-generation Gemini land on a Flash model (3.5 Flash, and now its successor 3.6 Flash). It is worth knowing what they have run into, both because it is your practical alternative and because one theme lines up suspiciously well with the reported reason Pro slipped. **Thinking tokens are the cost story.** They bill at the output rate, and developers report 3.5 Flash spending far more of them than a task warrants. The most carefully instrumented report logged a call with thinking level explicitly set to low that burned 63,853 thought tokens on a trivial prompt and returned zero output, fully billed. That is one very well-evidenced data point rather than a measured average, but the broader complaint recurs across several forum threads. **Flash is no longer the cheap tier.** At $1.50 in and $9.00 out per million tokens, 3.5 Flash costs exactly three times the Gemini 3 Flash Preview it succeeds ($0.50 / $3.00). If you want cheap, 3.1 Flash-Lite at $0.25 / $1.50 is the tier that now plays that role. One correction on a claim circulating in those threads: 3.5 Flash does not cost more than 3.1 Pro on list price, since 3.1 Pro Preview is $2.00 / $12.00. The argument people are actually making is about effective cost per completed task once thinking verbosity is counted, which is a different and harder-to-verify claim. Google has since shipped [Gemini 3.6 Flash and a new 3.5 Flash-Lite](https://geotoolbox.ai/blog/gemini-3-6-flash-vs-3-5-flash-lite), which reshuffle this low-cost lineup again. **Tool calling regressed.** Three independently filed breakages share the same shape, working on the older model and failing on 3.5 Flash: Google Search grounding combined with custom function declarations fails outright, with the model insisting it has no Search access; streaming with automatic function calling can end with empty content after the tool runs; and a contract change now requires clients to echo back a `thought_signature` from the tool call, returning a hard 400 to hand-rolled agent loops that do not. Google staff acknowledged the first with "looking into this" and no timeline. That last cluster is the interesting one. Reporting on Pro's delay says Google found structural failures in recursive tool calling and ordered a rebuild. Whatever is happening in Flash's tool handling appears to rhyme with it, which is at least a hint that the delay is about something real rather than schedule slippage. ## Should You Wait for Gemini 3.5 Pro? For most people, no. Reorganizing your stack around a model you cannot call, with specs Google has not published, is planning on rumor. Gemini 3.5 Flash is available today and already beats the previous Pro tier on coding and agentic work, which covers a lot of real use cases without waiting. Two things are worth holding in mind while you wait. First, cost. The reporting on the delay points at Pro's coding falling short of Google's internal bar, and if Pro lands anywhere near the rumored 10x pricing, the gap between "more capable" and "more expensive per task" will be the number that decides whether it is worth it. Model your budget on the real rate card when it ships, not the estimates. Second, do not build around the headline specs yet. A 2-million-token context window and a Deep Think mode would change how you design prompts and pipelines, but neither is confirmed. If you re-architect now and Google ships something different, you redo the work. The disciplined move is the one we took with [Grok 5](https://geotoolbox.ai/blog/grok-5): track what is confirmed, ignore what is guessed, and act when the model card exists. Anticipation is not availability. ## The Model Already Changed Under You Gemini 3.5 Pro is still a waiting game, but the model behind Google's AI answers changed the day Flash shipped. The useful question this week has little to do with which model to chase. It comes down to whether your pages get pulled into AI answers at all, on the surfaces that are already live. That is what geotoolbox checks. You can [run an AI readiness scan](https://geotoolbox.ai/tools/ai-readiness) to see whether the crawlers feeding Gemini and the other engines can actually reach and read your content, and whether the page carries the signals that make a citation more likely, before the next model lands and the gap widens. geotoolbox is the AI bot debugger for SEOs: it tells you why an engine is skipping your page, not just that it is. ## Frequently Asked Questions ### Is Gemini 3.5 Pro released yet? No. As of August 8, 2026, Gemini 3.5 Pro is still not generally available. Google announced the Gemini 3.5 family at I/O on May 19, 2026, but only Gemini 3.5 Flash shipped. On July 16, 2026, Bloomberg reported the rollout is months behind schedule and Alphabet shares fell about 4% that day; Google says only that it is currently testing Pro and has given no new date. ### When will Gemini 3.5 Pro come out? Google originally pointed to June 2026, then reporting pointed to a July 17, 2026 target. That date has now passed with no release. On July 16, 2026, Bloomberg reported Pro is months behind schedule, tied to coding performance that fell short of Google's internal bar, with no new date given. Earlier reporting also said Google rebuilt Pro from a new base model rather than the earlier 2.5 Pro one. Expect a limited preview first, not an immediate global launch. ### Does Gemini 3.5 Pro have a 2 million token context window? That figure is widely repeated but not confirmed by Google. The official announcement gives no context window for Pro at all, and the 1-million-token figure in circulation is associated with Gemini 3.5 Flash. Treat 2M as a rumor until Google publishes a model card. ### What is Deep Think in Gemini 3.5? Deep Think is described in third-party coverage as an extended-reasoning mode that spends more compute on hard problems. It is plausible and consistent with Google's direction, but it has not been attached to a published Gemini 3.5 Pro specification, so its exact behavior and availability are unconfirmed. ### Gemini 3.5 Pro vs Gemini 3.5 Flash, what's the difference? Flash is the fast, lower-cost tier; the current Flash model, 3.6 Flash, is the default in the Gemini app and AI Mode as of July 2026. Pro is the higher-ceiling tier and is not released. The notable confirmed fact is that 3.5 Flash already outperforms the previous Gemini 3.1 Pro on coding and agentic benchmarks. ### How much will Gemini 3.5 Pro cost? Unknown. Estimates range widely, from about $15 input and $60 output per million tokens down to roughly current Gemini 3.1 Pro pricing. Google has published no rate card, and the size of the disagreement is the clearest sign the real price is not public yet. --- ## What Is Grok Imagine? xAI's AI Image & Video Generator > Grok Imagine is xAI's AI image and video generator. Here's how to use it, whether it's free in 2026, what it can do, and how it stacks up against Sora and Veo. - Canonical: https://geotoolbox.ai/blog/grok-imagine - Published: 2026-06-28 · Updated: 2026-07-20 Grok Imagine is xAI's image and video generator, built into Grok. You type a prompt or upload a photo, and it returns a still image or a short video clip with sound. It is fast, it is opinionated about what it will and will not make, and the rules around price and access have changed more than once in 2026. Two things trip people up: whether it is still free, and what actually powers it. This guide answers both, walks through how to use it, and compares it fairly to Sora and Veo, current as of July 2026.
QuestionShort answer
What is it?xAI's text-to-image and image-to-video generator, inside Grok
Where do you use it?grok.com/imagine, the Grok iOS/Android app, or Grok on X
Current video modelGrok Imagine Video 1.5 (general availability June 16, 2026)
Is it free?Paid-first since early 2026. Free accounts get at most a thin image-only allowance; video, HD, and Spicy Mode need a paid plan
Video specsUp to 15 seconds, 480p or 720p, native audio in one pass, 24fps
Best forFast social, meme, and concept clips with sound, not 1080p broadcast work
## What Is Grok Imagine? Grok Imagine is the creative side of [Grok](https://geotoolbox.ai/blog/what-is-grok), xAI's chatbot. It does two jobs. It generates images from a text description, and it turns a still image into a short video, animating the scene and adding synchronized sound in the same pass. You reach it through an **Imagine** tab rather than the normal chat box. From there you can write a prompt, upload a reference photo to animate, edit an existing image, or chain several clips into a longer sequence. xAI also markets an Imagine Agent Mode that iterates on a prompt for you across a few steps instead of generating one shot at a time. The tool reportedly ships with four creative modes: **Normal**, **Fun**, **Custom**, and **Spicy**. Normal and Fun cover everyday generation, Custom gives you more control over style and motion, and Spicy is an age-gated mode for suggestive content, which we cover later because it carries real caveats. The headline feature is the video. Most AI video tools generate a silent clip and leave you to add audio afterward. Grok Imagine produces the picture and the sound together, dialogue, ambience, and effects timed to the action, so a clip comes out closer to finished. That single design choice is the thing the model is best known for, and it shapes most of the comparisons below. One framing worth getting straight up front: Grok Imagine is a feature inside Grok, not a separate product you sign up for on its own. Your Grok account, your plan, and your limits all carry over to it. ## Which Model Powers Grok Imagine? Aurora and Grok Imagine Video This is the part most guides get muddled, so here is the precise version. **Aurora** is xAI's in-house generative system. It launched first as the image model in [December 2024](https://en.wikipedia.org/wiki/Grok_(chatbot)), replacing the third-party Flux model Grok had borrowed until then. The video runs on a separate model, exposed in the xAI API as **grok-imagine-video** (the current release is `grok-imagine-video-1.5`). It is built on the same autoregressive approach Aurora pioneered, which is why people loosely say "Aurora powers the video." The cleaner way to put it: image and video are distinct models, and xAI brands the video one as grok-imagine-video. What makes Aurora unusual is that it is **autoregressive**, not diffusion-based. Tools like Sora and Runway generate a clip by denoising all the frames at once. Aurora generates each frame in sequence, with every new frame conditioned on the ones before it, the same next-step logic a language model uses to predict the next word. According to [Tech Times](https://www.techtimes.com/articles/318635/20260618/grok-imagine-video-15-goes-live-xai-tops-ai-video-leaderboard-86-percent-below-sora.htm), that design is why a camera move started in the first frame holds its trajectory through the last one, giving the model its stable motion and consistent subjects. The same choice explains a limitation we will come back to. Because the frames are generated one after another, pushing past 720p multiplies the work in a way the architecture cannot easily absorb, which is part of why 720p is the current ceiling. The native audio fits the same picture. Sound is generated inside that single forward pass, not bolted on later, so dialogue lands with lip-sync and effects match the on-screen action.
![How Grok Imagine makes a clip: a prompt or photo goes in, Aurora renders the image and grok-imagine-video animates it frame by frame, producing a short clip with native audio in one pass.](/blog/grok-imagine/grok-imagine-pipeline.png)
Aurora renders the still; grok-imagine-video animates it and generates the audio in the same pass.
## A Short History: How Grok Imagine Got Here Grok Imagine has moved fast, and the version names are easy to confuse. Here is the timeline that matters, as of July 2026.
DateWhat shipped
December 2024Aurora, xAI's in-house image model, debuts inside Grok
July 28, 2025Grok Imagine launches as a combined image and video tool
October 2025Version 0.9 arrives with faster, better video
January 28, 2026The Grok Imagine API opens to developers
February 1, 2026Imagine 1.0, with improved audio quality
June 16, 2026Grok Imagine Video 1.5 reaches general availability, plus a Fast variant
The current release is **Grok Imagine Video 1.5**. xAI first put it in [API preview on June 3](https://x.ai/news/grok-imagine-1-5), then moved it to general availability across the API, grok.com, and the mobile apps on June 16. A speed-tuned **Video 1.5 Fast** variant launched alongside it, generating a six-second 720p clip in roughly 25 seconds, down from 40 or more in the previous model. If you are reading an older guide that stops at version 0.9 or "Grok Imagine 1.0," assume its specs and prices are stale. This is a feature that has changed materially every couple of months. ## How to Access and Use Grok Imagine There are three official front doors: - **Web:** go to [grok.com/imagine](https://grok.com/imagine) and sign in with your Grok or X account - **Mobile:** open the Grok app on iOS or Android and tap the Imagine tab - **On X:** reach Grok from the X sidebar, where Imagine features depend on your plan and region A warning before you start. Search "Grok Imagine" and you will find sites like grokimagineai.net, grokvideo.ai, and dozens of similar names promising "free Grok Imagine." Those are third-party wrappers, not xAI. They run their own credit systems and want your sign-in. The real tool only lives at grok.com, in the official apps, and through the xAI API. To turn a photo into a video, the workflow is short: 1. Open Imagine and choose image or video 2. Upload a clear, well-lit photo, or generate one from a text prompt first 3. Describe the motion you want: a slow zoom, a pan, a character turning to speak 4. Pick a mode and a clip length, then generate 5. Use Extend from Frame to chain another clip onto the last one for a longer scene 6. Download the result as an MP4 Two practical habits make a real difference. Start from a sharp source image, since a soft or busy photo animates poorly. And change one thing per retry. The most common cause of a failed or off-target generation is asking for a complex scene plus several edits at once, so build it up step by step instead. One more thing if you are animating a person: expect the face to drift a little across frames. A sharp, front-lit source image and a shorter clip hold a likeness far better than a long, busy one. ## Is Grok Imagine Free? Plans, Limits, and Watermarks It depends on when you last checked, which is exactly why the question keeps coming up. The honest answer is a timeline, not a yes or no. Grok Imagine started effectively free. After the feature was used to mass-produce abusive images, xAI restricted image generation to paid subscribers in January 2026. By [March 20, 2026](https://www.calcalistech.com/ctechnews/article/rkbynj99bx), users were reporting that "Grok Imagine is no longer free for regular users," limited to X Premium subscribers and above, or SuperGrok, with xAI framing the change as a temporary technical measure. Through mid-2026 it stayed paid-first: free accounts get, at most, a thin image-only allowance inside Grok chat, while video, HD output, watermark-free results, and Spicy Mode all sit behind a paid plan. Here is the practical picture for consumers, drawn from our [Grok pricing breakdown](https://geotoolbox.ai/blog/grok-pricing):
PlanPriceWhat you get for Imagine
Free$0At most a thin, image-only allowance; watermarked, with tight reset windows
SuperGrok Lite~$10/moBasic image and short video generation
SuperGrok$30/moFull 720p video and image generation, higher limits, watermark-free
X Premium+~$40/moGrok Imagine access bundled with ad-free X
Limits are the part people complain about most, and they are a moving target. Reported caps have ranged from a handful of generations every couple of hours on the free tier to roughly 10 video generations in an eight-hour window on paid plans, with longer reset timers added over time. Treat any specific number you read, including ours, as a snapshot. xAI tunes these against server demand, so the figure can change without an announcement. On watermarks and rights: free output carries a corner watermark, and a paid plan removes it. Whether you can use a clip commercially is governed by xAI's terms, so check those before you publish a generated image or video for business use. ## Image and Video Quality: What You Actually Get Set expectations correctly and Grok Imagine is genuinely useful. Expect it to match Sora and you will be disappointed. For **images**, the xAI API exposes 1K and 2K outputs, which xAI's marketing rounds up to roughly four-megapixel. That is plenty for social posts, thumbnails, and concept art. For pure photorealism, character consistency, and 4K, dedicated image models still have an edge. For **video**, the [grok-imagine-video](https://docs.x.ai/developers/models/grok-imagine-video) model generates at 480p for drafts and 720p for final output, at a fixed 24 frames per second, in portrait, landscape, or square. A single generation runs up to 15 seconds. You can go longer by chaining clips with Extend from Frame, though community testing finds visible quality drift after two or three extensions, so longer sequences need care. The real limitation is resolution. 720p is the ceiling, and that shows up directly in the "it looks a bit mid" reaction you will see on Reddit. As covered earlier, this is architectural rather than a missing setting: Aurora's frame-by-frame design buys clean motion at the cost of cheap high-resolution scaling. xAI has said a higher-resolution Pro Mode is on the roadmap but has not given a date. The counterweight is the audio. Because sound is generated in the same pass as the video, a finished clip arrives with matched dialogue, effects, and ambience instead of a silent file you still have to score. For short-form work, that often matters more than the extra pixels. ## Content Moderation and "Spicy Mode" Grok markets itself as less filtered than its rivals, and Grok Imagine includes an age-gated **Spicy Mode** for suggestive content. That mode is also why much of the reporting on the tool is about its problems, so it is worth being precise. Spicy Mode is opt-in, restricted to paid plans, gated behind age verification, and mostly available in the app rather than on the web. Turning it on takes a paid plan, an age check, and switching on the sensitive-media toggles in your account settings. It is meant for suggestive or partial-nudity imagery of fictional characters. It does not permit sexual depictions of real, identifiable people, content involving minors, or non-consensual material, and many borderline prompts are blocked or blurred even with the mode on. The restrictions exist because of a serious episode. In late December 2025 and early January 2026, [Grok Imagine was used at scale](https://www.euronews.com/next/2026/01/05/grok-under-fire-for-generating-sexually-explicit-deepfakes-of-women-and-minors) to generate non-consensual sexual deepfakes of women, and reporting found it would also produce sexualized images of minors when prompted, including content involving a 14-year-old actress. The response varied by country. Indonesia and Malaysia went furthest, becoming the [first countries to block Grok outright](https://www.npr.org/2026/01/12/nx-s1-5674660/malaysia-indonesia-block-grok-ai-deepfakes) on January 11 and 12, 2026, while the European Union, the UK, France, and India opened investigations or scrutiny, and xAI faced lawsuits. xAI's own response was the paid-tier restriction covered above, plus tighter classifiers and an acceptable use policy that prohibits non-consensual intimate imagery and sexualized depictions of real people. The cases were still unfolding as of mid-2026. For everyday users, the practical residue of all this is the "content moderated, try a different idea" message. A second moderation pass runs after generation, and it sometimes flags innocent prompts because of a single trigger word or a resemblance to a real person. The fix is to rephrase and avoid naming or depicting real individuals, not to look for a way around the filter. Using AI to fabricate images of real people is exactly the kind of output that turns into a legal and reputational problem, in the same family of risks we describe in [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations): confident, plausible, and wrong in ways that land on whoever published it. ## Grok Imagine vs Sora 2, Veo, and Midjourney For video, Grok Imagine [currently tops the independent Image-to-Video Arena leaderboard](https://www.techtimes.com/articles/318635/20260618/grok-imagine-video-15-goes-live-xai-tops-ai-video-leaderboard-86-percent-below-sora.htm), ahead of Sora 2, Veo 3.1, Seedance 2.0, and Kling by user preference, though standings are volatile and its lead over models like Seedance 2.0 is narrow. The ranking is real, but read it carefully: the leaderboard measures average preference on general prompts, not fitness for a specific professional job. The clearer story is price and speed against quality.
ModelMax resolutionAPI price (per min)Native audioBest for
Grok Imagine 1.5720p~$4.20 (720p)Yes, in one passFast social and concept clips with sound
Sora 2 Pro1080p~$30 (1024p)YesHigher-fidelity work, where still available
Veo 3.11080p$9 to $24YesPolished 1080p output in the Google stack
MidjourneyImage-firstSubscriptionNoStylized, artful still images
The cost gap is the headline. At roughly [$4.20 per minute for 720p](https://gagadget.com/en/715382-grok-imagine-video-15-hits-number-one-and-its-86-cheaper-than-sora/), Grok Imagine undercuts Sora 2 Pro's $30 per minute by about 86 percent, and comes in well under Veo 3.1's $9 to $24. The picture shifted further in 2026 when OpenAI discontinued the Sora consumer app, leaving Grok and Google as the obvious choices for most creators. Google has since doubled down with [Gemini Omni](https://geotoolbox.ai/blog/gemini-omni), an any-input video model built around conversational editing. Where the competition wins back ground is resolution and pure fidelity. Sora 2 Pro and Veo output up to 1080p, which matters for client and broadcast work. For still images specifically, Midjourney remains the pick for artistic style and Google's models for photorealism, while Grok leans toward fast, atmospheric, wide-scene generation. We go deeper on the Google side in our [Grok vs Gemini comparison](https://geotoolbox.ai/blog/grok-vs-gemini). The bottom line: Grok Imagine wins on speed, native audio, and price, and trades away resolution to do it. Pick it for volume and iteration, not for the hero shot. ## Is Grok Imagine Safe to Use? For ordinary creative work, yes, with two cautions worth knowing. The first is data. Public posts on X can be used to train Grok by default, with an opt-out toggle in your settings, so think about that before you upload a personal photo to animate. The second is reputation. Given the deepfake history, uploading someone else's image to generate content of them is both a policy violation and a legal risk, and even your own AI-generated media should get a human review before you publish it under your name. The model can produce something that looks convincing and is subtly wrong, and the cost of that lands on the publisher, not the tool. ## What Grok Imagine Means for Your AI Visibility Here is the angle most creators miss. Grok Imagine sits inside Grok, and Grok is an answer engine that cites sources when it responds. When someone asks Grok, ChatGPT, or Google's AI about a tool, a product, or a brand, the engine pulls from a handful of pages it trusts. For this very topic, those pages are predictable. AI engines lean on YouTube, Reddit, and xAI's own properties when they explain Grok Imagine, which means the brands that get named are the ones publishing clear, current, well-structured information that a model can lift cleanly. In our experience helping brands with [AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility), the pattern holds across engines: the page that answers a question directly, with the facts an engine needs in a liftable form, is the page that gets cited. That is the same discipline behind getting picked up in [Google's AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo), and it is the work that turns a generative tool's popularity into visibility for you. If you want to see whether engines like Grok are citing your brand or your competitors, our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) tracks where you show up across the AI engines, so you can find the questions you are missing from and fix the pages that should be answering them. ## The Short Version Grok Imagine is one of the more capable, and more controversial, generative tools available, and it moves fast enough that any answer comes with a date attached. Use it for what it is good at, fast clips with sound, watch the changing limits and rules, and review anything before you put your name on it. ## Frequently Asked Questions ### Is Grok Imagine free in 2026? Mostly not. After image generation was restricted to paid plans in early 2026, free accounts kept at most a thin, watermarked, image-only allowance inside Grok chat. Video, watermark-free output, higher limits, and Spicy Mode all require a paid plan such as SuperGrok ($30/month) or X Premium+. ### How long can Grok Imagine videos be? A single generation runs up to 15 seconds at 480p or 720p. You can make longer clips by chaining generations with Extend from Frame, though quality tends to drift after two or three extensions. ### How do you turn a photo into a video with Grok Imagine? Open the Imagine tab, upload a clear photo, describe the motion you want, pick a mode and length, then generate. Sharp, well-lit images with simple backgrounds animate best, and changing one thing per attempt beats asking for a complex scene all at once. ### Is Grok Imagine better than Sora 2 or Veo? It depends on the job. Grok Imagine leads the image-to-video preference leaderboard and is far cheaper, around $4.20 per minute versus $30 for Sora 2 Pro, with native audio built in. But it caps at 720p, while Sora 2 Pro and Veo 3.1 reach 1080p, so for high-fidelity work the others still win. ### Can you use Grok Imagine images and videos commercially? Commercial use is governed by xAI's terms, and a paid plan removes the watermark. Check the current terms before publishing generated media for business, and never generate or publish content depicting real people without consent. ### Where do you find Grok Imagine? At grok.com/imagine on the web, in the official Grok app on iOS and Android, and through Grok on X. Third-party "free Grok Imagine" websites are not xAI and are best avoided. ## Sources - Grok (chatbot) - Wikipedia - `en.wikipedia.org/wiki/Grok_(chatbot)` - Grok Imagine 1.5 Preview - xAI - `x.ai/news/grok-imagine-1-5` - Grok Imagine API - xAI - `x.ai/news/grok-imagine-api` - grok-imagine-video model - xAI Docs - `docs.x.ai/developers/models/grok-imagine-video` - Grok Imagine Video 1.5 Goes Live - Tech Times - `techtimes.com/articles/318635/20260618/grok-imagine-video-15-goes-live-xai-tops-ai-video-leaderboard-86-percent-below-sora.htm` - Grok under fire for generating sexually explicit deepfakes - Euronews - `euronews.com/next/2026/01/05/grok-under-fire-for-generating-sexually-explicit-deepfakes-of-women-and-minors` - Malaysia, Indonesia become first to block Grok over AI deepfakes - NPR - `npr.org/2026/01/12/nx-s1-5674660/malaysia-indonesia-block-grok-ai-deepfakes` - Grok Imagine no longer free for regular users - Calcalist - `calcalistech.com/ctechnews/article/rkbynj99bx` - Grok Imagine Video 1.5 hits number one - gagadget - `gagadget.com/en/715382-grok-imagine-video-15-hits-number-one-and-its-86-cheaper-than-sora` --- ## Microsoft Copilot vs ChatGPT: Which Is Better? (2026) > Microsoft Copilot vs ChatGPT, honestly compared for 2026. Same OpenAI brain, different body: models, pricing, Office, privacy, and which one cites your brand. - Canonical: https://geotoolbox.ai/blog/microsoft-copilot-vs-chatgpt - Published: 2026-06-28 · Updated: 2026-08-07 The honest answer to Copilot vs ChatGPT starts with a fact worth stating plainly: they run on the same brain. The consumer Microsoft Copilot and ChatGPT both use OpenAI's GPT models, so on a quick question you would struggle to tell them apart. The difference is the body around that brain. ChatGPT is a standalone assistant you go to. Microsoft Copilot is wired into Word, Excel, Outlook, and your company's own files. Current as of 2026, this guide covers where that actually matters, the real prices and models, and one comparison most guides skip: which of the two is more likely to cite your brand when it answers. One note up front: the model names move almost monthly. Everything below is dated, and when a launch lands, the names change before the conclusions do. ## Copilot vs ChatGPT at a Glance Pick ChatGPT for standalone thinking, writing, coding, and the broadest tool ecosystem. Pick Microsoft Copilot for AI that works inside Microsoft 365 and can safely read your own emails, files, and meetings. On a one-off question, they are close enough to be interchangeable. The reason the choice is not obvious is that the two overlap more than either company admits. Both are multimodal chatbots, both search the web, both read files you upload, and both now run autonomous agents. The split shows up only when the work touches your organization's data or lives inside Office.
 Microsoft CopilotChatGPT (OpenAI)
MakerMicrosoftOpenAI
Underlying modelsOpenAI GPT-5.6 (Sol, Terra, Luna), plus Anthropic Claude and Microsoft's own MAI modelsOpenAI GPT-5.6 (Sol, Terra, Luna)
Reads your work filesYes, via Microsoft Graph (permission-scoped)Only files you upload, share, or connect
Acts inside Office appsYes, in Word, Excel, Outlook, TeamsNo
Web answers cite sourcesYes, on web answers, grounded in BingSometimes, when it decides to search
Best atOffice work, summaries, work-grounded answersWriting, reasoning, coding, ecosystem
Main paid planMicrosoft 365 Premium, $19.99/moPlus, $20/mo
The rest of this guide works through where those rows actually bite, starting with the question everyone asks first. ## Is Microsoft Copilot Just ChatGPT With Bing? Partly yes, and that is the honest version. The free Microsoft Copilot runs on OpenAI's GPT models, the same family as ChatGPT, and adds live web search through Bing, so for casual chat it really is close to ChatGPT with a Microsoft front end. That is why the "why pay extra" reaction is fair for the consumer version. But for the paid Microsoft 365 Copilot, the answer is no. Same brain, different body. The model reasoning is shared; everything wrapped around it is not. Microsoft 365 Copilot grounds itself in the [Microsoft Graph](https://geotoolbox.ai/blog/what-is-copilot), your organization's emails, documents, chats, and meetings, scoped to what you already have permission to see. ChatGPT cannot reach any of that on its own unless you connect or paste it in. There is also a model twist that breaks the "it's just ChatGPT" line. Copilot is no longer OpenAI-only. Microsoft has added Anthropic's [Claude](https://geotoolbox.ai/blog/what-is-claude-ai) models alongside GPT inside Copilot, and in June 2026 it started routing some work to its own in-house models too. So depending on the task and surface, Copilot may be running a model ChatGPT does not even offer. ChatGPT, by contrast, runs OpenAI models exclusively. And one naming trap worth clearing: **Microsoft Copilot is not GitHub Copilot.** They share a name and an owner and nothing else. GitHub Copilot writes code inside a developer's editor; Microsoft Copilot drafts your emails and summarizes your meetings. Paying for one does not give you the other, and most "Copilot vs ChatGPT" confusion online is really three products tangled together.
![Diagram showing Microsoft Copilot and ChatGPT both built on OpenAI's GPT family, with ChatGPT as a standalone app and Copilot wired into Microsoft 365, your files, and Bing.](/blog/microsoft-copilot-vs-chatgpt/same-brain-different-body.png)
Both run on OpenAI's GPT family; the difference is the body wired around it.
## Pricing: What Each One Actually Costs Both have a genuinely free tier, and for most personal use either is enough. You start paying for different reasons, which is what makes a straight price comparison misleading.
TierMicrosoft CopilotChatGPT (OpenAI)
Free$0 - chat, web answers with citations, image generation, voice$0 - GPT-5.6 (Luna), uncapped text chats from mid-August
EntryMicrosoft 365 Personal, $9.99/mo (Office apps + Copilot)Go, $8/mo
Main consumerMicrosoft 365 Premium, $19.99/moPlus, $20/mo
BusinessMicrosoft 365 Copilot, $30/user/mo enterprise; Business $18-$21 (up to 300 users, needs an M365 license)Business, about $25/user/mo ($20 annual)
Power / EnterpriseEnterprise pricing; Copilot Studio agents billed by creditsPro, $100-$200/mo; Enterprise custom
For a fair head-to-head, the closest match is **Microsoft 365 Premium at $19.99** against [**ChatGPT Plus at $20**](https://chatgpt.com/pricing). They are effectively tied on price, but you get different things. Premium puts Copilot inside your personal Word, Excel, and Outlook; Plus gives you OpenAI's full standalone toolset. Microsoft retired the old consumer Copilot Pro plan and [folded it into Microsoft 365 Premium](https://www.microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse/), so if a guide still lists "Copilot Pro" as the main consumer plan, it is out of date. The business comparison is where the sticker price misleads. Microsoft 365 Copilot is [$30 per user per month](https://www.microsoft.com/en-us/microsoft-365-copilot/pricing), but it is an add-on that sits on top of a qualifying Microsoft 365 subscription, so the all-in cost is higher than the headline number, often $45 to $90 per seat once you include the underlying license it requires. Our [Microsoft Copilot pricing guide](https://geotoolbox.ai/blog/copilot-pricing) breaks down every tier and the base-license math. Smaller businesses can get the Business version at $18 per user per month billed annually (a $21 list price). ChatGPT Business lands around $25 per seat, or about $20 billed annually, as a standalone product with no base license required. Is either worth paying for over free? For most casual users, no. Pay when you hit message caps often, want Copilot inside your Office apps, or need ChatGPT's heaviest models and tools. In our experience the teams that end up paying twice, Copilot for Office plus ChatGPT for everything else, are usually the ones who never sat down and matched the tool to the actual job. ## Which AI Models Each One Runs Here the shared-model point gets a wrinkle, and the version numbers actually matter. ChatGPT's flagship is **GPT-5.6**, which became generally available on July 9, 2026 across ChatGPT, Codex, and the OpenAI API, superseding GPT-5.5. It ships in three variants, Sol (flagship), Terra (balanced), and Luna (fast), and which one you get in a standard chat is set by your plan: since August 6, 2026 the free and Go tiers run Luna, while Plus and above run Sol, with Pro, Business, and Enterprise adding the Extra High effort level and Sol Pro. Terra is the odd one out and never appears in a standard conversation on any plan; you reach it in ChatGPT Work, in Codex, or through the API. Every ChatGPT tier runs an OpenAI model and nothing else. Copilot's default reasoning also runs on OpenAI's GPT models; as of July 9, 2026 Microsoft made **GPT-5.6** (Sol, Terra, Luna) the preferred model in Microsoft 365 Copilot, offered as a fast everyday mode and a slower "Think Deeper" reasoning mode, [the same flagship that powers ChatGPT](https://geotoolbox.ai/blog/how-does-chatgpt-work). The twist is that Microsoft no longer relies on OpenAI alone. Since [September 2025 it has offered Anthropic's Claude models](https://www.microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot/) (Sonnet 4 and Opus 4.1) inside Copilot Studio and its Researcher agent, its Copilot Cowork agent [launched on Anthropic's Opus 4.8 and Sonnet 4.6](https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/) (Sonnet 5 rolled in later in June 2026) and still uses those models, though Cowork is multi-model and GPT-5.6 became its preferred model on July 9, 2026. In June 2026 Microsoft also launched its own in-house MAI models that now handle some Copilot tasks. So the model picture splits two ways. On the shared GPT model, the underlying quality is comparable, because it is often the same model. But identical weights do not guarantee identical answers: Copilot's orchestration, system prompts, and tighter context limits can make a given reply feel thinner than raw ChatGPT, which is the kernel of truth behind the common "Copilot feels watered down" complaint. Copilot's real model-level edge is breadth, since it can route to Claude or Microsoft's own models, while ChatGPT stays inside OpenAI's lineup. One lasting frustration is that Copilot often does not tell you which model answered; Microsoft has started exposing model choice, but it remains more of a black box than ChatGPT. ## Microsoft 365 and Office Integration Office is where Copilot earns its price, and for a lot of buyers it settles the question. Copilot lives inside Word, Excel, PowerPoint, Outlook, and Teams, and on the Windows taskbar. It does not just answer in a side panel; it works inside the document you are in. Ask it to turn a draft into an executive summary and it can rewrite the file in place, rather than handing you text to paste back. The deeper win is grounding. With Microsoft 365 Copilot, you can ask it to "summarize the thread with the client" or "find the deck from last quarter," and it pulls from your real inbox, files, and meetings, scoped to your existing permissions. That is the gap a copy-paste workflow cannot close. ChatGPT answers this differently. It is not built into a single office suite, but it is far from an island, despite the common critique. Its connectors, Custom GPTs, and app integrations let it reach Google Drive, SharePoint, Slack, and thousands of apps through tools like Zapier, and it builds custom assistants that Copilot's consumer side cannot match. The difference is that ChatGPT mostly hands you output to move yourself, while Copilot does the work where the work already lives. The verdict here is clean. If your day runs on Microsoft 365, Copilot removes the copy-paste shuffle and that convenience is worth real money. If your stack is spread across many non-Microsoft tools, ChatGPT bends to more shapes. The same trade-off shows up across the field, and we break the wider pattern down in [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt). If Google's assistant is the real alternative you are weighing, see [Copilot vs Gemini](https://geotoolbox.ai/blog/copilot-vs-gemini). ## Web Search and Citations The two engines handle live information differently, and the difference matters more than it looks. When Copilot answers from the web, it **grounds in Bing by default**. It runs a search, reads the top pages, and writes an answer with clickable, numbered citations, the way [an AI search engine works](https://geotoolbox.ai/blog/how-does-ai-search-work). You can see where each claim came from and click through. (When it answers from your own files instead, it cites those, not Bing.) ChatGPT also searches the web, on by default, but it decides per query whether to look something up rather than grounding every answer. That makes it faster and cleaner on questions it already knows, but it is where ChatGPT can hand you a confident, well-written answer about something that changed last month without checking. Neither is reliable on freshness on its own. Independent hands-on testing has caught both pulling outdated sources for a "what happened this week" prompt, so the safe habit is the same for each: for anything recent or checkable, click the citations rather than trusting the summary. Both engines also [hallucinate](https://geotoolbox.ai/blog/ai-hallucinations), so treat either as a fast first draft of the truth, not the truth itself. That default-to-cite behavior is not just a usability detail. It is what turns this from a question about chat quality into one about whether anyone finds your brand. ## Your Data: Privacy and Security For business use, data handling often settles it, and it favors Copilot, with one real caveat. Microsoft 365 Copilot runs under **Enterprise Data Protection**: Microsoft says your work prompts and company data stay inside the Microsoft 365 service boundary and are not used to train the public models. It also inherits your existing file permissions rather than expanding them, so in theory Copilot only ever sees what you could already open. The caveat is real, though. As [security analysts have flagged](https://www.forcepoint.com/blog/insights/top-microsoft-copilot-security-risks), if a sensitive folder was accidentally shared too widely, Copilot makes it trivially easy to surface. It did not create the oversharing, but it strips away the obscurity that was hiding it, which is why a Copilot rollout often forces a long-overdue permissions cleanup. ChatGPT's defaults lean the other way for consumers. On the free and Plus plans it trains on your conversations by default unless you turn that off in settings, while its Business and Enterprise plans do not train on your data. So the practical rule splits by plan: consumer ChatGPT is the one to be careful pasting sensitive work into, and the question "what should I not tell ChatGPT" really means "do not paste anything you would not want retained." A last myth worth correcting: the free Copilot and Copilot Chat **can** read files you upload to them. What they cannot do is reach into your organization's data on their own, which only the paid Microsoft 365 Copilot does. If a guide says the free Copilot "can't read files," it is wrong. ## Writing, Coding, and Everyday Work On the actual tasks people run all day, the two trade wins, and the pattern is consistent enough to plan around. For **writing and creative work**, ChatGPT is the stronger pick. It varies its rhythm, picks up your tone, and follows complex creative instructions more reliably, which is why so many writers reach for it. Copilot tends to produce clean, structured, reliably on-brand drafts that are excellent for a business memo and a little flat for a creative campaign. Its stricter content filters also make it more likely to refuse or tone down opinion and fiction prompts that ChatGPT will simply write. For a deeper look at how the models differ on voice, see our [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) comparison. For **structured work**, Copilot earns its keep. In hands-on testing it tends to win on summarizing, pulling clean insights out of a spreadsheet, and generating quick visuals, often faster than ChatGPT. Its output is scannable and to the point, which is exactly what you want for a meeting recap or a data table. For **coding**, ChatGPT is better for explaining logic, debugging, and generating snippets in a chat. But if you mean writing code inside an editor, the tool you actually want is the separate GitHub Copilot, not Microsoft Copilot, which is the disambiguation that trips up half of these comparisons. User sentiment lines up with this split. Across G2 user reviews, ChatGPT rates slightly higher overall (around 4.7 versus 4.5), with its top marks for natural language and conversation quality, while Copilot's highest scores are for text summarization and in-app productivity. Both are strong; they are just strong at different things. ## Where Each One Falls Short Most comparisons skip this part. The topic skews negative in real-world use for a reason, so here are the weaknesses on both sides. **Copilot's quality is uneven**, and users feel it. It is genuinely good at summarizing meetings and drafting routine email, and weaker on complex reasoning and heavy spreadsheet work, where people still switch back to ChatGPT. A common complaint is that Copilot explains how to do a task in Office instead of just doing it, and reports of it slowing down or stalling on long sessions are frequent. It can also state wrong things confidently, like any [large language model](https://geotoolbox.ai/glossary/large-language-model). And many users find it unsolicited, pushed onto the Windows taskbar and into apps they did not ask for. Microsoft frames the "code red" and "retreat" headlines as a strategy revamp, not a shutdown, and Copilot is being expanded, not killed. None of this makes Copilot a bad tool; it makes it a specialized one, strongest inside the Office walls and weaker outside them. **ChatGPT's weakness is the flip side of its strength.** For work, it is an island: it does not know your company's files, calendar, or inbox unless you feed them in, which means a lot of copy-paste for anything grounded in your own data. Its per-query approach to web search can also miss something that changed recently, and on consumer plans it trains on your chats by default. None of these are dealbreakers, but they are the reasons a Microsoft-first team rarely makes ChatGPT its primary work assistant. ## The Part Nobody Compares: Which One Cites Your Brand Here is the comparison most "Copilot vs ChatGPT" guides skip, and the one that matters most if customers find you through AI. Both engines now answer questions by quoting web pages, but they do not pull those pages from the same place, so being cited by one is a weak predictor of being cited by the other. Copilot's web answers are grounded in the **Bing index**. On those public answers, if Bing cannot crawl and index your pages, Copilot will not cite you, and the playbook for fixing that is its own topic in [Copilot SEO](https://geotoolbox.ai/blog/copilot-seo). ChatGPT draws on a different mix, its own per-query search plus the sources its training favors, so the pages it surfaces for your category can look nothing like Copilot's list. In our experience auditing brands across AI engines, a company can be quoted confidently by ChatGPT for a question and be completely absent from Copilot's answer to the same question, simply because each engine trusts a different index. Most companies have no idea this gap exists, let alone which side they are missing from. The practical consequence is that optimizing for one engine does not carry to the other. You have to know where each one sources your category and earn a place in both, which starts with [measuring your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) per engine instead of guessing. ## Which Should You Use? (By Task) Skip the "it depends" and match the tool to the job.
If your main job is...UseWhy
Everyday questions and chatEitherSame GPT family; pick by free tier or ecosystem
Writing and creative workChatGPTWarmer voice, follows complex instructions better
Office docs, email, and meetingsCopilotEdits in place and summarizes from your real files
Work grounded in your own dataCopilotMicrosoft Graph access ChatGPT cannot match
Coding in a chatChatGPTBetter at explaining and debugging (in an editor: GitHub Copilot)
Current facts with citationsEitherCopilot cites Bing on web answers; click the sources on both
Standalone brainstorming and custom assistantsChatGPTCustom GPTs and the wider tool ecosystem
Sensitive company dataCopilotEnterprise Data Protection inside your tenant
Image generationEitherCopilot's Designer is fast; ChatGPT renders text in images better
Brand AI visibilityBothEach cites a different index; you need both
Most experienced users land in the same place: use both. Think and draft in ChatGPT because it reads more human, then do the work inside Office with Copilot because that is where the documents already live. With both free tiers, that hybrid costs nothing, and it sidesteps the false choice these comparisons usually force. ## The Comparison That Outlives the Model Roster The version numbers in this guide will be stale within months. What will not change as fast is the deeper split: Copilot and ChatGPT reason with the same OpenAI GPT family, but they live in different places and source their answers from different indexes. That decides where each one is useful, and whether your brand shows up when someone asks either one about your category. If customers increasingly find you through an AI answer instead of a blue link, the question stops being "which chatbot is better" and becomes "which one is recommending me, and which one has never heard of me." [Geotoolbox](https://geotoolbox.ai/features/citation-interceptor) tracks exactly that: which sources Microsoft Copilot's Bing-grounded answers, ChatGPT, and Google's AI Overviews cite for your category, and where your brand is missing from them, so you can earn a place in both engines instead of guessing. [See where you stand across engines](https://geotoolbox.ai/features/citation-interceptor) before the next model launch resets the board. ## Frequently Asked Questions ### Is Microsoft Copilot the same as ChatGPT? Not quite, though they overlap. The consumer Microsoft Copilot runs on the same OpenAI GPT family as ChatGPT, so for plain chat they are close. But Microsoft 365 Copilot adds grounding in your work email and files through the Microsoft Graph, enterprise data protection, the ability to act inside Office apps, and the option to run Anthropic's Claude models. For work that touches your own data, they are not the same tool. ### Is Copilot better than ChatGPT? Neither is better overall; it depends on the task. Copilot wins for work inside Microsoft 365 and answers grounded in your own files. ChatGPT wins for writing, standalone reasoning, coding in a chat, and the broader tool ecosystem. On everyday questions, they are close to interchangeable. ### Is Microsoft Copilot free, and is it better than free ChatGPT? Both have a free tier. The free Copilot covers chat, web answers with citations, image generation, and voice, and many reviewers give it a slight edge on free image generation. Free ChatGPT runs GPT-5.6 Luna, the fast variant, which replaced GPT-5.5 as its default on August 6, 2026. From the week of August 10 its text conversations are no longer rate-limited at all, though files, images, and voice keep their own caps, and the flagship Sol stays behind the paid tiers. On features they are close; the bigger difference is that free Copilot ties into the Microsoft apps and ChatGPT into a wider set of standalone tools. ### Does Microsoft Copilot use ChatGPT's models? Yes, in part. Copilot's default reasoning runs on OpenAI's GPT-5 series, the same family behind ChatGPT. But Microsoft also routes some Copilot tasks to Anthropic's Claude models, which ChatGPT does not offer, so they are no longer running identical model lineups. ### Why don't some people like Copilot? The common complaints are uneven quality compared with ChatGPT, a tendency to explain a task instead of doing it, occasional slowness, privacy worries about it surfacing overshared company files, and frustration that Microsoft pushed it into subscriptions and onto the Windows taskbar by default. ### Is Microsoft Copilot shutting down? No. Reports of a Copilot "code red" or "retreat" describe Microsoft reworking its AI strategy, not ending the product. Copilot is being expanded, including the 2026 move into autonomous agents like Copilot Cowork. ### Can I use Copilot and ChatGPT together? Yes, and many people do. A common pattern is to think and draft in ChatGPT, then do the work inside Office with Copilot. With both free tiers, running them side by side costs nothing. ## Sources - Microsoft 365 Copilot plans and pricing - enterprise and business pricing - `microsoft.com/en-us/microsoft-365-copilot/pricing` - Meet Microsoft 365 Premium - consumer plan replacing Copilot Pro, $19.99/month - `microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse` - Expanding model choice in Microsoft 365 Copilot - Anthropic Claude models in Copilot - `microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot` - Copilot Cowork is now generally available - 2026 agents on Anthropic models - `microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available` - Top Microsoft Copilot security risks, Forcepoint - data oversharing risk - `forcepoint.com/blog/insights/top-microsoft-copilot-security-risks` - ChatGPT plans and pricing - ChatGPT tiers and models - `chatgpt.com/pricing` - Microsoft Copilot, Wikipedia - timeline and underlying models - `en.wikipedia.org/wiki/Microsoft_Copilot` --- ## Perplexity Comet: What It Does, Costs, and Risks (2026) > Perplexity's AI browser is free, genuinely useful for research, and not safe to hand your inbox. What Comet does, what it costs, and how its agent hits your site. - Canonical: https://geotoolbox.ai/blog/perplexity-comet - Published: 2026-06-28 · Updated: 2026-08-10 Perplexity Comet is an AI browser that turns web browsing from something you do into something you delegate. Instead of opening tabs and clicking through pages yourself, you ask a built-in assistant to read, compare, and act for you. It is one of the most interesting products to come out of the AI race, and one of the most argued about. What follows is what it does, what it costs, whether it is safe to use, and what an agentic browser means for your site. Current as of August 2026. ## What Is Perplexity Comet? Perplexity Comet is a web browser built by Perplexity, the company behind the AI answer engine of the same name. It is built on [Chromium](https://en.wikipedia.org/wiki/Chromium_(web_browser)), the same open-source base as Google Chrome, so it looks and behaves like Chrome and imports your existing bookmarks and extensions. The difference sits in a sidebar: an AI assistant that can see the page you are on, read across your open tabs, and act on what it finds. That assistant is what makes Comet an "agentic" browser rather than a browser with a chatbot bolted on. A normal browser waits for you to click. Comet's assistant can do the clicking for you. The pitch is a shift from navigation to delegation, and we cover the broader pattern in our guide to [agentic AI](https://geotoolbox.ai/blog/agentic-ai). There is one point of confusion worth clearing up early, because people search for "Comet vs Perplexity" as if they are rivals. They are not. Comet is the browser. Perplexity is the [answer engine](https://geotoolbox.ai/blog/what-is-perplexity) that lives inside it, the same one you can use on the Perplexity website or app. Comet is Perplexity's attempt to own the whole browsing surface, not just the search box. Comet is also no longer the whole of that attempt. Per Perplexity's changelog, the company has been building out Perplexity Computer, an agentic workspace that sits above the browser and extends the same class of agent to documents, Microsoft 365 files, and, through a Personal Computer feature on Mac, local files and applications. Treat the specifics as a moving target, since Perplexity ships to this surface almost fortnightly, but the direction is the part that matters: Comet is becoming the browser surface of a broader agent platform rather than a standalone product. That reframes the security section below, because the question stops being only what an agent can do inside a tab. Comet first launched for Windows and macOS on July 9, 2025, then reached Android on November 20, 2025, and the iPhone on March 18, 2026. Perplexity's framing is direct: it wants to take browsing back from Chrome, with AI as the default way you interact with the web rather than an add-on. ## What Can Comet Actually Do? The core feature is the sidecar assistant, a panel that opens alongside whatever you are reading. You can ask it about the current page, have it summarize an article, a PDF, or a YouTube video, and ask questions that pull context from several open tabs at once. For research and reading, this is the part most people find genuinely useful. Beyond reading, Comet can act. Connect a Google account and the assistant gets read and write access to Gmail and Calendar, so it can answer questions about your schedule, dig through your inbox, or draft replies. Give it a higher-level instruction like "compare these three laptops" or "find a direct flight that morning," and it will visit sites, pull details, and work through the steps. It can also fill forms and build shopping carts. Here is the distinction the marketing skips. The free sidecar can run these tasks interactively, while you watch it work and step in when it gets stuck. What you pay for is autonomy. [Perplexity Pro](https://geotoolbox.ai/blog/perplexity-pricing) upgrades the agent's underlying model, and the Max plan adds Background Assistants, which grind through a to-do list on their own while you do something else. That has since become plural and schedulable: you can run several at once from a dashboard and set them on a recurring trigger, such as summarizing the morning's email and calendar at 8:30 each day. A Max-only Email Assistant shipped alongside them. And the agent is not reliable yet, though it is improving. Perplexity rebuilt the Comet Assistant in mid-2026 and shipped it to all users, claiming a 23% improvement over its predecessor on multi-step tasks along with longer-running jobs and the ability to work across tabs. That figure is Perplexity's own internal testing, not independent benchmarking, so treat it as a direction rather than a measurement. The field reports that set expectations before that rebuild are worth knowing anyway: one power user described a roughly 90% success rate on well-described tasks, which sounds fine until you realize that one in ten attempts fails, often halfway through, on anything involving logins, JavaScript-heavy pages, or precise form-filling. Reviewers who pushed harder reported failure rates closer to a third on delegated multi-step jobs. People also report Comet eating memory once the AI is active, and reviewers note there is little visible record of what the agent did, so a task that goes wrong is hard to debug. Treat it as a fast junior assistant you still have to watch, not an autopilot. ## Is Perplexity Comet Free? Yes, with a caveat. Comet was gated behind Perplexity's $200-per-month Max plan when it launched in July 2025, and the waitlist reportedly ran into the millions. That changed on October 2, 2025, when Perplexity [made Comet a free download worldwide](https://techcrunch.com/2025/10/02/perplexitys-comet-ai-browser-now-free-max-users-get-new-background-assistant/), with no subscription and no regional gate. So the "Comet costs $200 a month" line you may still see is out of date. The browser and the sidecar assistant are free. What the paid tiers add is the autonomous Background Assistant, higher usage limits, and Comet Plus, a $5-per-month bundle of content from publishers like CNN, The Washington Post, and Fortune that Perplexity pitches as an AI-flavored alternative to Apple News.
TierCostWhat you get
Comet (free)$0The browser plus the sidecar assistant: summarize, chat with tabs, basic tasks
Comet Plus$5/moPremium publisher content inside answers (also included with Pro and Max)
Perplexity Pro$20/moHigher limits, advanced models, Comet Plus included
Perplexity Max$200/moThe autonomous Background Assistant, top usage limits, priority models
Comet runs on Windows 10 and later, macOS Big Sur and later, Android 12 and later, and iOS 18 and later. You download it from perplexity.ai/comet or the mobile app stores, sign in with a Perplexity account, and import your bookmarks and extensions from Chrome. On the iPhone the app is a free download, with Pro and Max available as in-app purchases. ## Is Perplexity Comet Safe? This is where it gets uncomfortable. Comet's design creates a security problem that a normal browser does not have, and it has been demonstrated by more than one research team. The problem is indirect prompt injection. When you ask Comet to summarize a page, it feeds that page's content to its language model, and it does not reliably tell the difference between your instructions and instructions hidden inside the page. So a web page, an email, or a calendar invite can carry text that the assistant treats as a command. Because the agent acts inside your logged-in sessions, that command can do real damage. Brave's security team [showed this in August 2025](https://brave.com/blog/comet-prompt-injection/): they reported a working attack on July 25, Perplexity shipped a fix on July 27, the issue looked patched by August 13, and the details went public on August 20. Then [LayerX disclosed "CometJacking"](https://layerxsecurity.com/blog/cometjacking-how-one-click-can-turn-perplexitys-comet-ai-browser-against-you/), where a single crafted link tells the assistant to pull data from connected services like Gmail, encode it in base64 to slip past Perplexity's safeguards, and send it to an attacker. Perplexity initially marked that report as having no security impact. The attacks kept coming. [Trail of Bits](https://blog.trailofbits.com/2026/02/20/using-threat-modeling-and-prompt-injection-to-audit-comet/) published a threat-modeling audit, and in early 2026 [Zenity Labs demonstrated a version](https://zenity.io/research/pleasefix-vulnerabilities) triggered by a single calendar invite: once Comet processed the invite, the agent could read local files and pull credentials from an open 1Password session. Perplexity has since blocked that path with hard file-access boundaries. Note the scope: those boundaries constrain the browser agent. Perplexity Computer's Personal Computer feature deliberately connects an agent to your local files and apps as an opt-in surface with its own permissions, so it is a separate question rather than a hole in the same fence. Two common takes get this wrong. First, "it is just Chrome with a chatbot, so it is no riskier": the agent reading your tabs and acting in your sessions is a genuinely new attack surface. Second, "Perplexity said there was no security impact, so it is overblown" does not hold when six independent teams have demonstrated working prompt-injection exploits, several of them exfiltrating user data in proof-of-concept attacks. Guardio Labs' "Scamlexity" tests, for instance, tricked Comet into buying from a fake store, entering bank credentials on a phishing page, and obeying instructions hidden in a fake CAPTCHA to download a malicious file. Hacktron AI's is the clearest illustration of how wide the surface is: in August 2025 it chained a cross-site scripting flaw on a Perplexity subdomain with an over-permissive extension configuration to reach internal extension APIs, enabling arbitrary browser actions and reading Gmail contents. Perplexity hotfixed it within 24 hours and paid a $6,000 bounty. Perplexity has since added defenses, including a classifier that screens page content before the agent acts, and the rebuilt assistant now asks permission before it acts in your browser and holds that preference for the duration of a task. The disclosed exploits were patched as they surfaced. But indirect prompt injection is an unsolved class of problem, not a single bug, and new variants keep surfacing. One privacy point worth correcting: when AI features are on, the page content is processed on Perplexity's servers, not locally, despite some reviews implying otherwise. Until prompt injection is solved, do not run Comet's agent in a session where it has access to banking, health records, or work systems. Keep the connected accounts minimal, and review anything it does in a logged-in tab. Summarizing pages with no accounts connected is far lower risk than handing the agent your inbox. ## Perplexity Comet vs ChatGPT Atlas vs Chrome Comet is not the only AI browser anymore. OpenAI [launched ChatGPT Atlas](https://techcrunch.com/2025/10/21/openai-launches-an-ai-powered-browser-chatgpt-atlas/) on October 21, 2025, macOS first, with its agent mode tied to paid ChatGPT plans. Google has gone the other way, adding Gemini into Chrome rather than building a new browser around it. The cleanest way to think about Comet versus Atlas is what each one is built to do. Comet leans toward research: it is strong at reading a page, synthesizing across tabs, and showing you cited sources, which is a natural extension of Perplexity's answer engine. Atlas leans toward action: it is built around driving tasks and automating workflows. A useful shorthand is that Atlas is for automating and Comet is for annotating. Both share the same prompt-injection risk, because both let an AI act on web content.
 Perplexity CometChatGPT AtlasChrome
MakerPerplexityOpenAIGoogle
LaunchedJuly 9, 2025October 21, 20252008
Built onChromiumChromiumChromium
AI approachAI is the core; research and citation focusAI is the core; task and automation focusGemini added as an assistant
Agent costBackground Assistant on Max ($200/mo)Agent mode on paid ChatGPTLimited
If you live in Perplexity for research, Comet is the natural fit. If you mostly want an agent to grind through tasks, Atlas is built for it. Neither is ready to be your only browser, and you can read our take on [Perplexity against ChatGPT](https://geotoolbox.ai/blog/chatgpt-vs-perplexity) for the engine-level comparison. ## What Comet Means for Your Website If you run a site, the important thing about Comet is not whether you use it. It is that your visitors do, and their agent shows up in your data looking like a person. When someone delegates a task, the interactive agent loads your pages as ordinary Chromium, from the user's own device and inside their session. It does not announce itself the way Perplexity's declared crawlers do, so it does not show up as [PerplexityBot or the Perplexity-User agent](https://geotoolbox.ai/glossary/perplexity-user) that you can filter for in your logs. It looks like a normal visit. The result, as [HUMAN Security's analysis](https://www.humansecurity.com/ai-agent/perplexity-comet/) puts it, is traffic that comes from genuine human intent but executes at automated speed. The tell is behavioral: fast, dense, systematic navigation with few idle pauses, and a lot of it landing in your Direct or unattributed bucket in GA4.
![How a Comet agent visit reaches your site: user delegates, the agent loads your pages as Chromium, it looks human in analytics.](/blog/perplexity-comet/comet-agent-traffic-flow.png)
A Comet agent session loads your pages from the user's own browser, so it surfaces as ordinary traffic.
The instinct to block it is usually wrong, because behind the agent is a real customer trying to do something. The better move is to classify rather than block: let agents read and navigate, gate the state-changing actions like checkout or account changes, and watch the behavioral signals. The same logic is now playing out in [agentic commerce](https://geotoolbox.ai/blog/agentic-commerce), where Amazon escalated from a [November 2025 legal threat](https://techcrunch.com/2025/11/04/amazon-sends-legal-threats-to-perplexity-over-agentic-browsing/) to a lawsuit and won a March 2026 court order blocking Comet from shopping on Amazon, arguing the agent accessed its store without authorization. That injunction was stayed within weeks, and on August 4, 2026 the Ninth Circuit vacated it, holding that when a user tasks an agent to act on their behalf it is the user, not Perplexity, who "accesses" Amazon's computers, so Amazon is unlikely to prove a Computer Fraud and Abuse Act violation. The court stressed that its ruling is narrow and that this area of law "will doubtless change," and Amazon's trademark and state-law claims survive, so the dispute is not fully settled. It is the first appellate answer to whether an AI agent visiting a site on a user's behalf is authorized or unauthorized access, which makes it the single most consequential development for anyone running a website that agents will reach. In our experience, the brands that handle this well stop thinking about ranking alone and start thinking about being citable. Comet rewards content an assistant can read and lift cleanly into its sidebar: clear structure, answer-first passages, and pages that AI crawlers can actually reach. That is the same discipline behind [optimizing for Perplexity](https://geotoolbox.ai/blog/perplexity-seo) and [tracking whether AI engines mention you](https://geotoolbox.ai/blog/how-to-track-ai-visibility). ## Should You Use Perplexity Comet? If you are curious about where browsing is heading, Comet is worth a look, and the free sidecar assistant for summarizing and tab research is the lowest-risk way in. If you mostly read and research, you will get value from day one. If you need reliability, or you spend your day in banking, healthcare, or anything with sensitive logged-in sessions, wait. The agent is not dependable enough to trust unsupervised, and the security questions are real and unresolved. There is no harm in keeping Chrome as your daily driver and trying Comet on the side. The bigger shift is the one underneath all of this. As assistants start reading and acting on the web for people, the question stops being where you rank and becomes whether your content is the thing they read, trust, and quote. ## Frequently Asked Questions ### What is Perplexity's Comet? Comet is an AI browser made by Perplexity, built on Chromium. It puts an AI assistant in a sidebar that can summarize the page you are on, answer questions across your open tabs, and carry out tasks like comparing products or filling forms. It launched in July 2025 and became a free download in October 2025. ### Is Comet from Perplexity free? Yes. The browser and the sidecar assistant are free with no subscription, since October 2, 2025. Paid Perplexity plans add extras: the autonomous Background Assistant on the $200-per-month Max plan, and Comet Plus ($5 per month) for premium publisher content. ### Is Perplexity the same as Comet? No. Perplexity is the AI answer engine, available on its website and app. Comet is Perplexity's web browser, with that same engine and an assistant built into it. You can use Perplexity without Comet, but Comet is built around Perplexity. ### Is Perplexity Comet safe to use? For low-stakes reading and summarizing, the risk is modest. For agentic tasks, be careful. Researchers at Brave, LayerX, Trail of Bits, Zenity, Guardio, and Hacktron have all demonstrated prompt-injection or extension-level attacks where hidden instructions on a page hijack the assistant. Avoid running the agent in sessions with banking, health, or work data until the issue is better solved. ### What is the difference between Perplexity Comet and ChatGPT Atlas? Both are Chromium-based AI browsers. Comet leans toward research and citation, fitting Perplexity's answer-engine roots. ChatGPT Atlas, from OpenAI, leans toward automating tasks. A shorthand: Atlas is for automating, Comet is for annotating. Both carry the same prompt-injection risk. ### What devices and platforms support Comet? Comet runs on Windows 10 and later, macOS Big Sur and later, Android 12 and later, and iOS 18 and later. It launched on Windows and macOS on July 9, 2025, Android on November 20, 2025, and iPhone on March 18, 2026. ## The Takeaway Comet is a real preview of where browsing is going, not a finished product. It is free, it is genuinely useful for research, and it is not safe to hand your inbox to yet. Whether or not you adopt it, your visitors will, and their agents are already reading your site. That changes the job. The pages that win the agentic web are the ones an assistant can reach, read, and quote without friction. You can [check whether AI agents and crawlers can actually read your site](https://geotoolbox.ai/tools/ai-readiness) in a couple of minutes. geotoolbox is the AI bot debugger for SEOs, built to show you exactly what Comet, ChatGPT, and the rest see when they look at your pages. ## Sources - Comet (browser) - Wikipedia (engine, platform launch dates) - `en.wikipedia.org/wiki/Comet_(browser)` - Perplexity's Comet AI browser now free; Max users get new Background Assistant - TechCrunch - `techcrunch.com/2025/10/02/perplexitys-comet-ai-browser-now-free-max-users-get-new-background-assistant` - Agentic browser security: indirect prompt injection in Perplexity Comet - Brave - `brave.com/blog/comet-prompt-injection` - CometJacking: one click can turn Comet against you - LayerX - `layerxsecurity.com/blog/cometjacking-how-one-click-can-turn-perplexitys-comet-ai-browser-against-you` - Using threat modeling and prompt injection to audit Comet - Trail of Bits - `blog.trailofbits.com/2026/02/20/using-threat-modeling-and-prompt-injection-to-audit-comet` - PleaseFix: zero-click agent hijacking in Comet - Zenity Labs - `zenity.io/research/pleasefix-vulnerabilities` - What is Perplexity Comet and why is it on my website? - HUMAN Security - `humansecurity.com/ai-agent/perplexity-comet` - Amazon sends legal threats to Perplexity over agentic browsing - TechCrunch - `techcrunch.com/2025/11/04/amazon-sends-legal-threats-to-perplexity-over-agentic-browsing` - One-click UXSS in Perplexity Comet - Hacktron AI - `hacktron.ai/blog/perplexity-comet-uxss` - Amazon.com Services LLC v. Perplexity AI, Inc. (docket) - CourtListener - `courtlistener.com/docket/71874820/amazoncom-services-llc-v-perplexity-ai-inc/` - AI Agents and the CFAA: Amazon.com Services v. Perplexity AI - Volokh Conspiracy - `reason.com/volokh/2026/06/19/ai-agents-and-the-cfaa-amazon-com-services-v-perplexity-ai/` - Ninth Circuit rules on AI agent access to third-party websites under the CFAA (injunction vacated, Aug 4 2026) - Cooley - `cooley.com/news/insight/2026/2026-08-06-ninth-circuit-rules-on-ai-agent-access-to-third-party-websites-under-cfaa` - OpenAI launches an AI-powered browser: ChatGPT Atlas - TechCrunch - `techcrunch.com/2025/10/21/openai-launches-an-ai-powered-browser-chatgpt-atlas` --- ## Perplexity Pricing 2026: Plans, Pro, Max, and Is It Worth It? > Perplexity's 2026 pricing explained: Free, Pro ($20), Max ($200), the $10 student plan, the credit system, and an honest take on whether Pro is worth it. - Canonical: https://geotoolbox.ai/blog/perplexity-pricing - Published: 2026-06-28 · Updated: 2026-08-11 Perplexity has six plans: a free tier, Pro at $20 a month, Max at $200, a $10 student plan, and two enterprise tiers. Most people only ever need the free plan or Pro. The hard part is not the headline prices, it is what changed underneath them in 2026: a credit system, an agent called Perplexity Computer, and quietly tightened limits. Here is what each plan costs, what you get for it, and who should pay for which, as at the update date on this page. ## Perplexity Pricing at a Glance
PlanPriceBest forHeadline of what you get
Free$0Casual lookupsThe default Sonar model with citations, around three Pro Searches a day
Pro$20/mo or $200/yrDaily researchModel switching, near-unlimited Pro Search, file uploads, image generation, Comet Plus, Computer access
Max$200/mo or $2,000/yrHeavy agent and automation usersEverything in Pro plus 10,000 monthly Computer credits, Model Council, unlimited research
Education Pro$10/moVerified students and educatorsThe Pro plan at half price, verified through SheerID
Enterprise Pro$40/seat/mo or $400/yrTeamsPro plus SSO, admin controls, internal knowledge search, no training on your data
Enterprise Max$325/seat/mo (or custom)Large orgs running agentsEnterprise Pro plus 15,000 monthly credits and the highest limits
One caveat on the numbers. The source of truth is [Perplexity's own plans page](https://www.perplexity.ai/pro), but it sits behind a bot wall, and the third-party guides that rank for "perplexity pricing" disagree with each other on the fine print: daily search caps, whether the API credit still ships, whether Enterprise Max is a fixed price or custom. Every number above is cross-checked against Perplexity's live plans page and multiple current sources, and where they genuinely conflict, this guide flags the conflict instead of presenting one figure as settled. The short version: the free plan is better than most people expect, Pro is the best value in the lineup if you do real research more than a couple of times a week, and Max is a niche tool for people running automated agent work, not a "better answers" upgrade. ## Is Perplexity Free? What You Get Without Paying Yes, and the free plan is genuinely usable. You get unlimited basic searches on [Perplexity](https://geotoolbox.ai/blog/what-is-perplexity)'s own Sonar model, every answer comes with citations, and you get a small daily allowance of Pro Searches, the deeper multi-step searches that read several sources before answering. Perplexity's own help center now lists three a day, though some older guides still say five. Either way, it resets daily and it is enough for occasional use. What you give up on free is real but specific. You cannot switch models, so you are stuck with Perplexity's fast default rather than picking GPT, Claude, or Gemini for a given task. Image generation is off. File uploads are limited. Deep Research, the long report-style mode, is capped to roughly one run a month. And you do not get access to Perplexity Computer, the agent layer that the paid plans are now built around. For someone who asks Perplexity a handful of questions a week, the free plan covers it. The moment you start leaning on it for work, you will hit the Pro Search ceiling and the single-model limit fast, and that is the upgrade Perplexity is counting on. ## Perplexity Pro ($20/Month): What You Actually Get Pro is $20 a month, or $200 a year if you pay annually, which works out to about $16.67 a month. It is the plan most paying users are on, and the jump from free is the biggest value-per-dollar step in the lineup. The headline upgrade is model switching. On Pro you choose which frontier model answers a given question, drawing from the current OpenAI, Anthropic [Claude](https://geotoolbox.ai/blog/what-is-claude-ai), and Google Gemini releases plus Perplexity's own Sonar. The exact model names move almost monthly, so treat any specific version you read as a snapshot rather than a contract. You also get near-unlimited Pro Search, Deep Research at a much higher allowance (commonly cited at around 20 runs a day), the Labs report-builder, large file uploads, AI image generation, and access to premium data sources like financial and academic databases. Two things matter more than the feature list. First, Pro includes access to Perplexity Computer plus a one-time bundle of around 4,000 Computer credits, not a monthly refill. That distinction trips people up, and the next section explains why. Second, Pro has historically included a small monthly credit toward Perplexity's developer API, though some 2026 reporting says consumer plans no longer bundle it, so confirm it at signup rather than counting on it. The word "unlimited" comes with asterisks now. Through 2026 Perplexity quietly tightened limits on its most demanding features. Users [reported](https://www.androidauthority.com/perplexity-pro-advanced-ai-limits-reduced-3667942/) hitting weekly caps on the heaviest models after as few as three to five queries a day, file-upload caps triggering after two uploads, and per-response token limits cut from 200 to 100. Perplexity said the changes mostly affected accounts tied to promotional codes, citing fraud and resale. Two things are worth knowing before you pay anyway: the heaviest models are rationed more tightly than "unlimited" suggests, and some users report being quietly served Perplexity's cheaper fallback model without choosing it. Pro is still a high-ceiling plan for almost anyone, but check that you are getting the model you picked. If you are deciding between monthly and annual, the annual plan saves you about $40 a year. Given how fast Perplexity changes its plans, paying monthly for the first couple of months and switching to annual once you know you will keep it is the lower-risk play. ## Perplexity Max ($200/Month): Who It's Really For Max is $200 a month, or $2,000 a year. It is ten times the price of Pro, and the instinct is to assume that buys ten times better answers. For everyday questions, it does not. Max mostly buys scale and automation. What you get over Pro is volume. Max includes 10,000 Perplexity Computer credits a month (plus a one-time bonus when you upgrade), unlimited Labs and Research, priority access to the newest models, enhanced media generation, and Model Council, a feature that runs the same question through several models at once and synthesizes the answers. The base models that answer your Pro queries also answer your Max queries, so for a normal single-model question the answer quality is the same. Max's one genuine quality lever is Model Council on hard problems; the rest of the gap is how much agent and research work you can throw at it before you hit a wall. The hidden detail that makes Max make sense for the right person is the spending cap. Every Max account has a monthly credit spending limit, defaulting to $200 and adjustable up to $2,000. When you hit it, running tasks pause. That tells you exactly who Max is for: someone whose agent workflows would otherwise blow through credits, who wants a predictable ceiling on heavy automated use. For everyone else, Max is the wrong plan. If your day is asking questions and reading cited answers, you are unlikely to use enough of what it offers to justify it, and Pro gives you the same model quality for a tenth of the cost. ## The Credit System Explained (Perplexity Computer) This is the part most pricing guides skate over, and the part that surprises people when their first bill arrives. In early 2026 Perplexity launched Perplexity Computer, an agent that orchestrates around 19 different models to carry out multi-step jobs on your behalf, the kind of "go do this whole task" work we cover in our guide to [agentic AI](https://geotoolbox.ai/blog/agentic-ai). Computer does not run on your normal search allowance. It runs on credits. A credit is metered agent compute. Browsing a page, filling a form, running a step in a workflow, all of it draws down a credit balance. The problem is that Perplexity does not publish a clear "this action costs this many credits" table, so the meter is hard to predict. One user reported a single 40-minute task burning roughly 23,000 credits, more than a Max plan's entire monthly allowance. Worse, a task that fails partway still consumes the credits it used, and they are not automatically refunded.
![Perplexity Computer credits by plan as of June 2026: Free has none, Pro gets a one-time 4,000-credit bonus, Max gets 10,000 a month, Enterprise Pro 500 a month per seat, and Enterprise Max 15,000 a month per seat.](/blog/perplexity-pricing/perplexity-computer-credits-by-plan.png)
Pro's credits are a one-time bonus; only Max and the enterprise tiers refill monthly.
Here is how the allowances stack up in practice:
PlanComputer creditsWhat that signals
FreeNoneNo Computer access
Pro~4,000, one-time bonusEnough to try it, not to live on it
Max10,000 per monthBuilt for ongoing agent use
Enterprise Pro~500 per month per seatLight, shared team use
Enterprise Max15,000 per month per seatHeavy team automation
Once Pro's bonus credits are gone, serious Computer use means moving to Max or buying more. That is not an accident of plan design. When the Financial Times reported that Perplexity's annual recurring revenue [jumped past $450 million](https://ca.finance.yahoo.com/news/perplexity-arr-tops-450m-pricing-132500539.html) in early 2026, it tied the surge directly to the shift toward AI agents and usage-based pricing. The credit economy is the business model, which is exactly why it pays to understand the meter before you switch it on. ## Student, Education and Free-Pro Deals: What's Real in 2026 Search "how to get Perplexity Pro for free" and you will find a graveyard of expired offers presented as if they still work. Here is what is actually live versus dead as of mid-2026. The reliable discount is **Education Pro at $10 a month**, the full Pro plan at half price for students and educators who verify through SheerID. It has dipped to around $5 during promotional windows. Perplexity has also run a referral program giving verified students free months that stack, reported as valid through May 31, 2026, so by now treat it as expired unless Perplexity has renewed it. Separately, US government and military workers with a qualifying email address have been offered a free year of Pro at various points. The carrier and fintech deals are where people get burned. The widely shared **PayPal and Venmo free-Pro offers expired on December 31, 2025**. The Airtel promotion in India ran out in January 2026. New regional partnerships appear and lapse constantly, so by the time a "free Perplexity Pro" post is circulating, the window has often already closed. Before you sign up through any third-party "free Pro" link, check the expiry date and whether it is region-locked. Several 2026 offers advertised as no-card trials later required a card and auto-renewed. If a deal asks for payment details for something billed as free, treat that as the catch. ## Enterprise Pro vs Enterprise Max ($40 vs $325 per Seat) For teams, Perplexity sells two tiers. Enterprise Pro is $40 per seat per month, or $400 a year per seat. On top of everything in individual Pro, it adds single sign-on, SCIM provisioning, role-based admin controls, SOC 2 Type II compliance, and Internal Knowledge Search, which lets a team query an uploaded shared file repository. Crucially, company data on enterprise plans is not used to train Perplexity's models. Each seat comes with a modest monthly Computer credit allowance, commonly cited around 500. Enterprise Max is the heavy tier at $325 per seat per month, or $3,250 a year, though the most recent guides note it can also be quoted as custom pricing for the largest deployments. It raises every limit and bumps each seat to roughly 15,000 Computer credits a month, aimed at teams running real agent workloads rather than just team search. Verified educational institutions have also been offered Enterprise Pro at a reduced per-seat rate. For most teams, Enterprise Pro is the sensible default. Enterprise Max only earns its price when seats are genuinely running automated Computer tasks at volume. ## Perplexity API and Sonar Pricing (For Developers) If you are building on Perplexity rather than chatting with it, the [Sonar API](https://docs.perplexity.ai/docs/getting-started/pricing) is billed separately from any consumer subscription, on tokens plus a per-request search fee.
ModelInput / output (per 1M tokens)Notes
Sonar$1 / $1Lightweight search-grounded answers
Sonar Pro$3 / $15Deeper, multi-step search
Sonar Reasoning Pro$2 / $8Reasoning over search results
Sonar Deep Research$2 / $8Plus citation, reasoning, and search-query fees
On top of tokens, each model charges a per-1,000-request fee that scales with how much search context you pull, roughly $5 to $14 for the standard models and up to about $22 for the agentic Pro Search mode. The exact figures move, so price a real workload against the live docs before you commit. This is a developer concern, not a consumer one, and it is separate from the credits that power Perplexity Computer inside the app. ## Perplexity vs ChatGPT Plus, Claude Pro and Gemini: Do You Need to Pay? At $20 a month, Perplexity Pro sits in the same price bracket as the other major AI subscriptions, and the most common real question is not "is Pro good" but "do I need it if I already pay for one of these."
PlanPriceBest at
Perplexity Pro$20/moLive, cited web research across multiple models in one place
ChatGPT Plus$20/moGeneral assistant work, writing, coding, image and voice
Claude Pro$20/moLong-document reasoning and careful writing
Google AI Pro~$20/moGemini tied into Google Search, Workspace, and Android
Perplexity's distinct edge is answer-engine search: every response is grounded in live sources you can click, and you can pick the model behind it. If your day is research where the citation matters, that is worth paying for on its own, and it is a different job from a general chatbot. We go deeper on the head-to-head in [Perplexity vs ChatGPT](https://geotoolbox.ai/blog/chatgpt-vs-perplexity). The real answer to "do I need both" is usually no. If you already pay for ChatGPT Plus or [Claude Pro](https://geotoolbox.ai/blog/claude-pricing) mainly to chat, write, or code, Perplexity overlaps enough that a second $20 is hard to justify, and the free Perplexity plan covers the occasional cited lookup. If cited, multi-model research is the core of your work, Perplexity earns its place even alongside another subscription. The same logic applies if you are weighing [Gemini's plans](https://geotoolbox.ai/blog/gemini-pricing): pick the tool whose main job matches yours, and resist paying twice for the overlap. ## Is Perplexity Pro Worth It? An Honest Verdict For the right user, Pro at $20 is one of the easier yes calls in AI subscriptions. For the wrong one, it is $240 a year you will barely touch. It comes down to one thing: how often you do cited research. Pro is worth it if you run real research more than a couple of times a week. Marketers, analysts, researchers, journalists, consultants, and students who need answers with sources they can check get the most out of it: live web results, real citations, and the freedom to pick the best model for a task, all in one place. If you qualify for the $10 student rate or a free year, the math gets even easier. It is not worth it if your AI use is casual, if your work is mostly creative writing or coding where a general assistant fits better, or if you already pay for [ChatGPT Plus](https://geotoolbox.ai/blog/chatgpt-pricing) or Claude Pro and do not specifically need cited search. The free plan handles light use well enough that paying adds little. On Max, the verdict is narrower. In our experience, Max only makes sense for a small group: people running automated agent workflows or Model Council comparisons at volume who want a predictable spending ceiling. For everyone else it is paying ten times the price for capacity you will not use, with no gain in answer quality. If you are asking whether you need Max, you almost certainly do not. ## Where Perplexity's Paid Plans Fall Short Paying Perplexity does not buy you out of its real weaknesses, and a pricing guide that skips them is not worth much. The biggest one is that **price does not buy accuracy**. A Tow Center study covered by the [Columbia Journalism Review](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php) found AI search tools got citations wrong in over 60% of news queries, with Perplexity around 37%, and notably the paid variants were not reliably cleaner than the free ones. The failure mode is the dangerous kind: a real URL attached to a claim the source never made. Pro and Max give you more searches and better models, not freedom from confident, well-cited errors. You still have to verify. The credit system is the second issue. Because per-task credit costs are not published, Max users have been genuinely surprised by how fast credits vanish, and a failed task still burns them with no refund. Third, several users report billing friction: auto-renewal is the default, the annual refund window is short, and deleting your account does not cancel your subscription. None of this makes Perplexity a bad product. It makes it one you should sign up for with eyes open, especially before stepping up to Max. ## How to Cancel Perplexity and Whether You Get a Refund Cancelling is simple once you know it will not happen by itself. On the web, open Settings, find Subscription under Perplexity Pro, click Manage Subscription, then Cancel. If you subscribed through the iPhone or Android app, you have to cancel in the App Store or Google Play instead, because the app store owns that billing. Cancelling stops the next renewal and you keep Pro until the end of the period you already paid for. One trap to avoid: deleting your Perplexity account does not cancel the subscription, so people who just remove the app keep getting charged. Refunds are harder. Perplexity generally does not refund unused time, though consumer law widens the window in some regions: monthly plans are often refundable within about 24 hours and annual plans within about 72 hours, while EU, UK, and Turkey customers get 14 days. If support stonewalls a charge you believe is wrong, a card-issuer chargeback is the realistic fallback. The cleaner habit is the one from earlier: stay monthly until you are sure, and only move to annual once you know you will keep it. ## Getting Cited by Perplexity Is the Other Half of the Equation Everything above is about what you pay Perplexity. If you run a brand or publish content, there is a more valuable question hiding underneath it: when Perplexity answers questions about your market, does it cite you? That is a different game from a subscription. Perplexity sends referral traffic and visibility to the sources it cites in answers, and it has even started paying some of them through [Comet Plus](https://www.searchenginejournal.com/perplexity-launches-comet-plus-shares-revenue-with-publishers/554596/), a publisher program that shares revenue on an 80/20 split favoring publishers. Being in the cited set is the visibility play, and it is one you earn through how your content is structured and sourced, not through a $20 plan. We break down the tactics in our guide to [getting cited in Perplexity](https://geotoolbox.ai/blog/perplexity-seo). This is what we built [geotoolbox](https://geotoolbox.ai) for. Our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) finds the offsite sources that Perplexity and the other answer engines cite in your space where your brand is missing, grades each one by how strong the evidence is, and ranks them by how often you are left out. Paying for Pro makes you a better Perplexity user. Getting cited gives Perplexity a reason to surface you. If you are spending on AI subscriptions anyway, the second one is where the return is. ## Frequently Asked Questions ### Is Perplexity cheaper than ChatGPT? They cost the same at the entry tier: both Perplexity Pro and ChatGPT Plus are $20 a month. Perplexity is cheaper for students at $10, and it has a free plan that is more research-capable than ChatGPT's free tier for cited search. At the top end, Perplexity Max and ChatGPT's $200 tier are also matched. Price is rarely the deciding factor between them; the job you need done is. ### Is a Perplexity subscription worth it? For people who do cited research regularly, yes, Pro at $20 is strong value. For casual users, the free plan is enough, and for anyone already paying for another AI assistant they mainly chat with, a second subscription is hard to justify. Match the plan to how often you actually need sourced answers. ### How much does Perplexity Pro cost? Perplexity Pro is $20 a month, or $200 a year if you pay annually (about $16.67 a month). Verified students and educators pay $10 through Education Pro. There is no permanent free trial of Pro, though Perplexity has run short trials and free-month promotions at various points, so the practical way to test it is to pay for one month and cancel if it is not for you. ### Why is Perplexity so expensive? Among the individual plans, only Max is expensive, at $200 a month, and it is priced for heavy agent and automation use rather than ordinary search (enterprise seats run higher still, up to $325). The cost reflects Perplexity routing your queries to frontier models it licenses from OpenAI, Anthropic, and Google, plus the credit-metered compute behind Perplexity Computer. For normal use, the $20 Pro plan, or free, is the relevant price. ### What is the downside of Perplexity? Paying does not eliminate inaccuracy. Independent testing found Perplexity, like all AI search tools, still misattributes sources a meaningful share of the time, sometimes pairing a real link with a claim it does not support. The credit system is also hard to predict, and some users report billing and cancellation friction. Treat it as a fast research aid you still fact-check, not an oracle. ### Is Perplexity Max worth $200 a month? For almost everyone, no. Max does not produce better answers than Pro; it produces more capacity, more Computer credits, and Model Council comparisons. It is worth it only if you run automated agent workflows or high-volume research and want a predictable spending ceiling. If you are unsure whether you need it, you do not. ### Can I get Perplexity Pro for free? Sometimes, legitimately. Verified students and educators get Pro for $10 through SheerID, and free-month referral and government-email offers have run during 2026. Be skeptical of carrier and fintech "free Pro" deals shared online, though: the major PayPal and Venmo offers expired at the end of 2025, and many circulating links are out of date or region-locked. ### Why is Perplexity being sued? The lawsuits are about how Perplexity gathers and uses content, not about its pricing. Several news publishers have sued it for copyright infringement: News Corp's Dow Jones and the New York Post filed first, in October 2024, the New York Times followed in December 2025 after an earlier cease-and-desist, and CNN sued in May 2026. Perplexity disputes the claims, and the cases are still working through the courts. It is worth knowing as context, but it does not change what your subscription costs or what the plans include. - [Perplexity Sonar API pricing](https://docs.perplexity.ai/docs/getting-started/pricing) (docs.perplexity.ai) - [Perplexity's Comet browser is now free for everyone](https://www.theverge.com/news/790419/perplexity-comet-available-everyone-free) (The Verge) - [Comet free, Max gets the Background Assistant](https://techcrunch.com/2025/10/02/perplexitys-comet-ai-browser-now-free-max-users-get-new-background-assistant/) (TechCrunch) - [AI Search Has a Citation Problem](https://www.cjr.org/tow_center/we-compared-eight-ai-search-engines-theyre-all-bad-at-citing-news.php) (Columbia Journalism Review, Tow Center) - [AI search engines are confidently wrong](https://fortune.com/2025/03/18/ai-search-engines-confidently-wrong-citing-sources-columbia-study/) (Fortune) - [Perplexity ARR tops $450M after pricing shift](https://ca.finance.yahoo.com/news/perplexity-arr-tops-450m-pricing-132500539.html) (Financial Times via Yahoo Finance) - [The New York Times sues Perplexity over copyright](https://www.cnbc.com/2025/12/05/the-new-york-times-perplexity-copyright.html) (CNBC) - [CNN sues Perplexity over alleged AI copyright theft](https://www.cnn.com/2026/05/28/media/cnn-sues-perplexity-ai-copyright) (CNN Business) --- ## What Is Microsoft Copilot? Versions, Pricing & Models (2026) > What is Microsoft Copilot? A plain guide to all the Copilot products, what they cost, how it works, Copilot vs ChatGPT, and how to show up when it answers. - Canonical: https://geotoolbox.ai/blog/what-is-copilot - Published: 2026-06-28 · Updated: 2026-08-13 Microsoft Copilot is the AI assistant Microsoft has put almost everywhere: a free chatbot on the web, a button in Word and Excel, an assistant built into Windows, and a paid layer that reads your work email. If you want the plain version of what Copilot actually is, which versions exist, what they cost, and whether it is just ChatGPT with a Microsoft badge, this is it, current as of August 2026. First, the hard part. "Copilot" is not one product. Microsoft uses the name for at least seven different things across four price models, and its own pages rarely stop to say which one you are looking at. Worse, **GitHub Copilot is a separate product** that a lot of people mean when they say "Copilot." This guide is about Microsoft's AI assistant. We will sort the family out, then get to the part most explainers skip: because Copilot now answers questions by quoting web pages, what it says about your company is something you can influence. ## What Is Microsoft Copilot? **Microsoft Copilot is an umbrella brand for a family of AI assistants built on large language models, wired into Microsoft's apps, your work data, and live web search.** Microsoft's own definition is "a conversational, AI-powered assistant" that helps you write, summarize, analyze, code, and automate tasks. The key word is *family*. There is no single app called "Copilot" that does everything; there is a consumer chatbot, a version inside Microsoft 365, one built into Windows, a security tool, and a platform for building custom ones. Under the hood, every Copilot combines three things: a language model that understands your prompt and writes the answer, a source of grounding (the live web, or your own files and email when you are signed in to a work account), and an integration layer that lets it act inside an app, like drafting an email in Outlook or building slides in PowerPoint. If the product feels new, the engine is not. Copilot started life as **Bing Chat**, which [Microsoft launched on February 7, 2023](https://en.wikipedia.org/wiki/Microsoft_Copilot). Microsoft began rebranding it to "Microsoft Copilot" on September 21, 2023, and folded its various AI features under that one name. Like [Google's Gemini](https://geotoolbox.ai/blog/what-is-gemini), it is a general-purpose assistant that the maker has pushed into its entire ecosystem rather than a single standalone app. ## Which Copilot? The Products Hiding Under One Name The names are where the confusion starts, so here is the whole family in one place. When someone says "Copilot," they could mean any of these:
ProductWho it's forWhat it doesPrice (2026)
Free CopilotAnyoneWeb and app chatbot: questions, writing, image generation, voiceFree
Microsoft 365 Premium (replaced consumer Copilot Pro)Individuals and familiesOffice apps plus Copilot inside Word, Excel, PowerPoint, Outlook$19.99/month
Microsoft 365 CopilotBusinessesCopilot grounded in your work email, files, chats, and meetings$30/user/month (Enterprise); $21 list for Business, up to 300 users
Copilot ChatWork and schoolSecure, web-grounded chat using your work identity (no automatic access to your org's data)Free with eligible Microsoft 365 plans
Copilot StudioBuilders and ITLow-code tool to build and govern custom Copilot agents~$200/pack (25,000 credits) or $0.01/credit pay-as-you-go
Windows CopilotWindows 11 usersSystem assistant: change settings, open apps, search, chatFree in eligible Windows editions
Security CopilotSecurity teamsAI assistant for threat hunting and incident responseMetered by Security Compute Units
GitHub CopilotSoftware developersSeparate product: AI coding assistant inside code editorsFree tier, then Pro and Business plans
![The Microsoft Copilot family: one brand split into Consumer, Microsoft 365, Platform and OS, and Developer products, with GitHub Copilot flagged as separate.](/blog/what-is-copilot/copilot-family-tree.png)
"Copilot" is one brand spanning consumer, Microsoft 365, platform, and developer products, with GitHub Copilot run separately.
Two distinctions cause most of the confusion. The first is **GitHub Copilot versus Microsoft Copilot**. They share a name and an owner and nothing else. GitHub Copilot writes code inside a developer's editor; Microsoft Copilot drafts your emails. Paying for one does not give you the other. The second is **free Copilot versus Copilot Chat versus paid Microsoft 365 Copilot**. Free Copilot and Copilot Chat can both search the web, and you can upload a file for either to read. What they cannot do is reach into your organization's data on their own. Only the paid Microsoft 365 Copilot grounds itself in your whole work corpus, summarizing your Teams meetings and drafting replies from your real inbox without you pasting anything in. That gap, between an assistant that knows the public web and one that knows *your* work, is the whole reason the business version costs money. One more naming note worth keeping straight: Microsoft retired the standalone consumer **Copilot Pro** subscription and folded its features into **Microsoft 365 Premium**, [which it introduced on October 1, 2025](https://www.microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse/) at $19.99 a month. If a guide still talks about Copilot Pro as the main consumer plan, it is out of date. ## How Microsoft Copilot Works Copilot answers in one of two modes, and knowing which one you are in explains almost everything about what it can and cannot do. In **web mode**, which is what the free Copilot and Copilot Chat use, it works like any [AI search engine](https://geotoolbox.ai/blog/how-does-ai-search-work): it runs a search, reads the top web pages, and writes an answer that quotes them with clickable, numbered citations. It knows the public internet and whatever you upload to it, but it does not reach into your organization's files on its own. In **work mode**, which only the paid Microsoft 365 Copilot has, it adds a second source: the **Microsoft Graph**. That is Microsoft's index of your organization's content, your emails, documents, chats, calendar, and meetings, scoped to what you already have permission to see. This is what lets it say "summarize the thread with the client" or "find the deck from last quarter." It is also why the business version raises real privacy questions, which we will get to. The models behind it shift often. As of July 2026, Copilot's default reasoning runs on **OpenAI's GPT-5.6**. Microsoft made it the preferred model for Microsoft 365 Copilot across Word, Excel, PowerPoint, Chat, and Cowork when GPT-5.6 became generally available on July 9, 2026. It runs as a fast everyday mode and a slower "Think Deeper" reasoning mode, the same family that powers ChatGPT. But Microsoft has been adding choice: since [September 2025 it offers Anthropic's Claude models](https://www.microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot/) (Sonnet 4 and Opus 4.1) as options inside Copilot Studio and its Researcher agent, and its autonomous features can still route to Anthropic models. Copilot is also becoming more of a multi-model router than a pure-GPT one: [Office Watch, reporting Bloomberg, July 2026](https://office-watch.com/2026/microsoft-mai-models-excel-outlook/) reports that Microsoft has quietly begun routing some high-volume Excel and Outlook prompts to its own in-house "MAI" models to cut costs, even with GPT-5.6 as the headline model. So "what model is Copilot" no longer has one answer; it depends on the surface and the task. Copilot is less a single AI than a router that sits between you, a choice of frontier models, and either the web or your work data, with Microsoft's apps as the place it does the work. ## Is Microsoft Copilot Free? What It Actually Costs Yes, there is a genuinely free Copilot, and for most personal use it is enough. The free version at copilot.microsoft.com and in the mobile apps gives you chat, web answers with citations, image generation, and voice, with usage limits during busy periods. For everyday use, you start paying when you want one of two things. If you want Copilot **inside your personal Office apps**, that comes with Microsoft 365 Premium at $19.99 a month. If you want Copilot to **read your company's work**, that is Microsoft 365 Copilot at [$30 per user per month](https://www.microsoft.com/en-us/microsoft-365-copilot/pricing) for enterprises. Smaller businesses can get the Business version (up to 300 users) at a $21 list price, though Microsoft has been running a 15% promotion that brings it to $18 per user per month through December 31, 2026 (month-to-month, with no annual commitment, is $25.20). For every tier broken down, including the base-license math and the Studio and Cowork credit pricing, see our [Microsoft Copilot pricing guide](https://geotoolbox.ai/blog/copilot-pricing). Don't confuse the free Copilot Chat that shows up at work with the paid version. It looks identical but cannot reach your emails and files on its own. If a colleague says "Copilot summarized my inbox" and yours will not, the difference is a $30 license, not a setting. ## Microsoft Copilot vs ChatGPT The most common question about Copilot is whether it is just ChatGPT wearing a Microsoft badge. It is a fair question, because the consumer Copilot does run on the same OpenAI models, and it grew out of the same partnership. The real differences are not the model; they are the wiring around it.
DimensionMicrosoft 365 CopilotChatGPT
Underlying modelsOpenAI GPT-5.6 (plus Claude options and in-house MAI models)OpenAI GPT-5.6
Reads your work filesYes, via Microsoft Graph (permission-scoped)Only files you share or connect
Acts inside Office appsYes, in Word, Excel, Outlook, TeamsNo
Your data trains public modelsNo, under Enterprise Data ProtectionDepends on your plan and settings
Best forWork grounded in your company's dataStandalone chat, brainstorming, coding
Microsoft 365 Copilot is grounded in the **Microsoft Graph**, so it can answer using your own emails, files, and meetings, which ChatGPT cannot reach unless you paste them in. For businesses, Copilot also runs under **Enterprise Data Protection**: Microsoft keeps your work prompts and company data inside the Microsoft 365 service boundary and does not use them to train the public models. And Copilot can *act* inside Word, Excel, and Outlook, while ChatGPT mostly hands you text to copy out. For a standalone chatbot to brainstorm or write, ChatGPT and free Copilot are close, and many people find ChatGPT a little sharper. For an assistant that already lives inside the Microsoft tools your company pays for and can safely read your work, Copilot is doing something ChatGPT is not built to do. We go deeper on this in [Microsoft Copilot vs ChatGPT](https://geotoolbox.ai/blog/microsoft-copilot-vs-chatgpt), and the same split shows up across the field in [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt). ## What's New in 2026: Agents and Copilot Cowork Through 2026, Microsoft's whole pitch for Copilot shifted from "AI that answers questions" to **AI that does multi-step work on its own**. The clearest example is **Copilot Cowork**, [which became generally available on June 16, 2026](https://www.microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available/). Cowork can take a long, multi-step job, run it in the background, and let you check progress from your phone while it works. Notably, it launched on Anthropic models (Opus 4.8 and Sonnet 4.6, with Claude Sonnet 5 rolling into Cowork on July 2, 2026) and can still use them, though Cowork is multi-model: GPT-5.6 became its preferred model on July 9, 2026, alongside the rest of Microsoft 365 Copilot. On top of a Microsoft 365 Copilot license, Microsoft bills it through pay-as-you-go "Copilot Credits" at about $0.01 per credit rather than a flat fee. This is part of a broader move into [agentic AI](https://geotoolbox.ai/blog/agentic-ai): assistants that take actions, not just give answers. [Copilot Studio](https://geotoolbox.ai/blog/copilot-studio) lets companies build their own agents that reach into CRM and database systems, and Microsoft has added prebuilt ones like Researcher and Analyst. If you tried Copilot a year ago and found it was a glorified chatbot, the 2026 version is a different kind of tool, for better and worse. ## Where Copilot Falls Short Microsoft's own pages skip this part. Across user forums and reviews, the same complaints come up, and they are worth knowing before you rely on it. **Quality is uneven.** Copilot is strong at summarizing meetings and drafting routine email, and weaker at complex reasoning and spreadsheet work, where users still reach for ChatGPT or Claude. Like any [large language model](https://geotoolbox.ai/glossary/large-language-model), it can state wrong things confidently, so its output needs checking. **Privacy is the real catch for business.** Copilot inherits your existing file permissions; it does not expand them. But that cuts both ways: [as security analysts have flagged](https://www.forcepoint.com/blog/insights/top-microsoft-copilot-security-risks), if a sensitive SharePoint folder was accidentally shared too widely, Copilot makes it trivially easy to find. Copilot did not create the oversharing, but it removes the obscurity that was hiding it. **It feels forced.** Microsoft has pushed Copilot into Microsoft 365 subscriptions and onto the Windows taskbar, and a vocal share of users actively look for how to turn it off. Headlines about a Copilot "code red" or "retreat" describe a strategy revamp, not a shutdown; Copilot is not going away, but the rollout has been contentious. ## Copilot Is Also an Answer Engine: How to Show Up When It Answers If you run a website, this is the part that matters most. When someone asks the free Copilot or Copilot Chat a question, it does not invent the answer alone. It searches the web, reads pages, and **cites them with clickable footnotes**. Those citations are real traffic and real authority, and they go to specific pages. The question for any brand is whether one of them is yours. This makes Copilot an [answer engine](https://geotoolbox.ai/blog/what-is-answer-engine-optimization), not just a chatbot, and it is a visibility channel you can influence the same way you influence Google. [Copilot SEO](https://geotoolbox.ai/blog/copilot-seo) is the full playbook; in short, three things decide whether Copilot can cite you: - **Reachability.** Copilot's web answers are grounded in Bing's search index, so being indexable by Bing is the entry ticket. If Bing and the other [AI crawlers cannot fetch your pages](https://geotoolbox.ai/blog/ai-crawlers), you are invisible before the contest even starts. - **Liftable answers.** Copilot quotes self-contained passages that directly answer a question. Pages that bury the answer in fluff get skipped in favor of ones that state it plainly. - **Corroboration.** Like most AI engines, Copilot favors claims that show up consistently across several trusted sources, not a single unverified page. In our experience auditing brands across AI engines, most companies have no idea whether Copilot can even see them, let alone whether it cites them. That blind spot is the whole reason we built geotoolbox: to show you where the AI assistants people now ask are sourcing their answers, and whether your pages are in the running. ## Frequently Asked Questions ### Is Microsoft Copilot the same as ChatGPT? No, though they overlap. Consumer Copilot runs on the same OpenAI models as ChatGPT, but Microsoft 365 Copilot adds grounding in your work email and files via the Microsoft Graph, enterprise data protection, and the ability to act inside Office apps. For plain chat they are similar; for work that touches your own data they are not. ### Is Microsoft Copilot free? Yes. The free Copilot on the web and in the mobile apps covers chat, web answers, image generation, and voice. You pay only for Copilot inside your Office apps (Microsoft 365 Premium, $19.99/month) or for Copilot that reads your company's work (Microsoft 365 Copilot, $30/user/month). ### Is GitHub Copilot the same as Microsoft Copilot? No. GitHub Copilot is a separate product, an AI coding assistant for software developers that works inside code editors. Microsoft Copilot is the general assistant for writing, search, and Office work. A subscription to one does not include the other. ### Why don't some people like Copilot? The common complaints are uneven quality compared to ChatGPT, privacy concerns about it surfacing overshared company files, and frustration that Microsoft pushed it into subscriptions and onto the Windows taskbar by default. ### Is Copilot shutting down? No. Reports of a Copilot "code red" or "retreat" describe Microsoft reworking its AI strategy, not ending the product. Copilot is being expanded, including the 2026 move into autonomous agents. ### Does Copilot cite its sources? Yes. In its web-grounded mode, Copilot returns answers with inline, clickable citations to the pages it used, which is why being one of those cited pages is worth optimizing for. ## Which Copilot Is Right for You? If you just want a free AI assistant, use the free Copilot or compare it with ChatGPT and [Gemini](https://geotoolbox.ai/blog/what-is-gemini); our [Copilot vs Gemini](https://geotoolbox.ai/blog/copilot-vs-gemini) breakdown shows which fits your work. If you live in Office and want AI in your documents, Microsoft 365 Premium is the consumer path. If you run a business and want AI that safely works across your email, files, and meetings, Microsoft 365 Copilot is the one that earns its price. And if your audience is starting to ask AI assistants instead of typing into a search box, the more pressing question is not which Copilot to buy, but whether Copilot can find and cite you when it answers. You can [check your AI readiness](https://geotoolbox.ai/tools/ai-readiness) and see exactly where you stand. ## Sources - Microsoft Copilot, Wikipedia - launch and rebrand timeline, underlying models - `en.wikipedia.org/wiki/Microsoft_Copilot` - Microsoft 365 Copilot plans and pricing - enterprise and business pricing, Copilot Chat - `microsoft.com/en-us/microsoft-365-copilot/pricing` - Meet Microsoft 365 Premium - consumer plan replacing Copilot Pro, $19.99/month - `microsoft.com/en-us/microsoft-365/blog/2025/10/01/meet-microsoft-365-premium-your-ai-and-productivity-powerhouse` - Expanding model choice in Microsoft 365 Copilot - Anthropic Claude models in Copilot - `microsoft.com/en-us/microsoft-365/blog/2025/09/24/expanding-model-choice-in-microsoft-365-copilot` - Copilot Cowork is now generally available - 2026 agent capabilities, Copilot Credits (GA on Opus 4.8 and Sonnet 4.6) - `microsoft.com/en-us/microsoft-365/blog/2026/06/16/copilot-cowork-is-now-generally-available` - Anthropic's Claude Sonnet 5 in Microsoft 365 Copilot - Sonnet 5 rolled into Copilot/Cowork after the June 16 GA - `techcommunity.microsoft.com/blog/microsoft365copilotblog/available-today-anthropics-claude-sonnet-5-in-microsoft-365-copilot/4532188` - Top Microsoft Copilot security risks, Forcepoint - data oversharing risk - `forcepoint.com/blog/insights/top-microsoft-copilot-security-risks` --- ## What Is Perplexity? The AI Answer Engine, Explained > What is Perplexity AI? A current guide to the answer engine: who owns it, Comet, pricing, Perplexity vs ChatGPT, accuracy, and what it means for your brand. - Canonical: https://geotoolbox.ai/blog/what-is-perplexity - Published: 2026-06-28 · Updated: 2026-07-22 Perplexity is the AI that answers your question with a written, cited answer instead of a list of links. It is the tool people increasingly open before Google when they want a fast, sourced answer rather than ten tabs to read through, and it is often the top pick among [the best AI search engines](https://geotoolbox.ai/blog/best-ai-search-engines) for cited research. Most "what is Perplexity" explainers stop at "it is an AI search engine." That is true and not very useful. This covers what actually matters: who builds it, what it costs, whether you can trust it, and what it means for whether AI mentions your brand, current as of June 2026.
![A Perplexity-style answer card with numbered inline citations linking to cited source pages.](/blog/what-is-perplexity/perplexity-cited-answer-anatomy.png)
Perplexity's defining trait: a written answer where every claim carries a numbered, clickable citation.
## What Is Perplexity? **Perplexity is an AI-powered answer engine: it runs a live web search for almost any question, then uses a large language model to write a direct, conversational answer with the sources cited inline as numbered footnotes.** You ask in plain language, and instead of blue links you get a short written answer with clickable citations under each claim. That makes it a different kind of tool from the two things people confuse it with. A traditional search engine hands you a ranked list of links and leaves the reading to you. A general chatbot like ChatGPT writes from what its model learned during training, which can be out of date and has no sources attached. Perplexity sits in between: it behaves like a research librarian that goes and finds current pages, then summarizes them and shows you where each line came from. It is built around its citations, which is the single most important thing to understand about it.
Search engine (Google)Chatbot (ChatGPT)Answer engine (Perplexity)
What you getA ranked list of linksA written answer from training dataA written answer from live sources
Sources shownThe links are the resultOften none by defaultInline citations on each claim
FreshnessMixed; can rank old pagesLimited by training cutoffLive web search on most queries
Best atNavigation, shopping, mapsWriting, brainstorming, codeResearch and fact-finding with sources
Perplexity is what you reach for when you want a sourced answer to a real question, fast. It is one of the clearest examples of an [answer engine](https://geotoolbox.ai/glossary/answer-engine), the category that is reshaping how people find information online. ## Perplexity the Product vs Perplexity the Metric One quick source of confusion worth clearing up. In machine learning, **"perplexity" is a long-standing metric** that measures how well a [language model](https://geotoolbox.ai/glossary/large-language-model) predicts text: lower perplexity means the model is less "surprised" by the next word, so a lower score is better. The company named itself after that concept, but the product and the metric are separate things. If you came here looking for the statistics term, that is the other one. For the rest of this article, Perplexity means the answer engine. ## Who Owns Perplexity, and Is It Legit? **Perplexity AI, Inc. is an independent, privately held US company based in San Francisco, founded in August 2022** by Aravind Srinivas (CEO), Denis Yarats (CTO), Johnny Ho, and Andy Konwinski. The founders came out of AI research and engineering roles at places like OpenAI, Google DeepMind, Meta, and Databricks. It is not owned by Google, Microsoft, or OpenAI, and despite a common assumption, it is not a Chinese app. It is also well funded by names you will recognize. According to its [Wikipedia profile](https://en.wikipedia.org/wiki/Perplexity_AI), investors include Nvidia, Jeff Bezos, Databricks, SoftBank, and NEA, and in December 2025 Cristiano Ronaldo took a stake and a brand partnership. The valuation has climbed fast: past $1 billion in April 2024, to $14 billion by mid-2025, and roughly $21 billion in early 2026 after a later funding round. So the "is this a sketchy startup" worry is misplaced, and yes, Jeff Bezos really did invest. The scale is real too. At Bloomberg's Tech Summit in 2025, Srinivas said Perplexity handled [780 million queries in May 2025](https://techcrunch.com/2025/06/05/perplexity-received-780-million-queries-last-month-ceo-says/), growing more than 20% month over month. Its reach jumped further when India's Airtel handed [free Perplexity Pro](https://www.airtel.in/press-release/07-2025/airtel-partners-with-perplexity-powers-every-single-of-its-360mn-customers-with-perplexity-pro/) to its 360 million customers in mid-2025, one of several moves that pushed adoption well beyond the US. **Can you buy Perplexity stock?** Not directly. Perplexity is privately held, so there is no public ticker to buy on a stock exchange. The only ways to hold a position are private or secondary markets that most people cannot easily access, or buying shares of its public investors. If you see a "Perplexity stock" listing, treat it with suspicion, because the company itself is not publicly traded. ## How Perplexity Works Under the hood, Perplexity is a retrieval-augmented answer engine. When you ask a question, it reads your query, runs a live web search, pulls the most relevant passages from the pages it finds, and then hands those to a language model that writes the answer and attaches a citation to each claim. That pattern of fetching real sources first and generating the answer second is called [retrieval-augmented generation](https://geotoolbox.ai/glossary/retrieval-augmented-generation), and it is why the answers come with footnotes instead of arriving from nowhere. If you want the longer version, we wrote a full breakdown of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work). Two design choices shape the experience. The first is **model choice**. Perplexity is not locked to a single model. It builds its own Sonar models, and on the paid tier it lets you pick the engine that writes the answer, choosing among frontier models from OpenAI's [GPT](https://geotoolbox.ai/blog/seo-for-chatgpt), Anthropic's [Claude](https://geotoolbox.ai/blog/what-is-claude-ai), and Google's [Gemini](https://geotoolbox.ai/blog/what-is-gemini). This is the part people get wrong most often: Perplexity is not built by OpenAI and is not "ChatGPT with sources." It is its own product that can route your question to several different models. The second is **focus modes**. Instead of searching the whole web every time, you can point Perplexity at a specific kind of source: Academic for peer-reviewed papers, Social for forum and community discussion, Finance for market data, and so on. Narrowing the source pool is how you get a research-grade answer instead of a generic one. ## What Perplexity Can Do Perplexity started as a search box and has turned into a small family of products in 2026. You can use it free on the web at perplexity.ai, in the iOS and Android apps, or inside its own browser. The core is still the answer engine, but the surfaces around it are where the recent action is. The everyday workhorse is **search with Pro Search and Deep Research**. A normal query returns a quick cited answer; Pro Search asks clarifying questions and digs deeper, and Deep Research runs many searches in a row and writes up a longer report. **Pages** lets you turn a research thread into a clean, shareable article. There are also Shopping and Finance surfaces that pull live product and market data into the answer. Image generation is available on the paid tiers, but treat it as a convenience rather than a reason to pick Perplexity over a dedicated image tool. The bigger 2026 story is the move beyond a website. **Comet** is Perplexity's own [web browser](https://geotoolbox.ai/blog/perplexity-comet), built on Chromium and launched in July 2025, then opened up as a free download in October 2025. It puts an AI assistant inside the browser that can summarize the page you are on, compare what is across your open tabs, and carry out multi-step tasks like filling forms or drafting an email. Alongside it, **Perplexity Assistant** and the newer **Computer** product push into agentic territory, where you describe an outcome and the tool does the clicking. In practice, Perplexity's durable edge is the cited, live answer and the browser-and-agent push around it. The image generation and creative features are competitive but not clearly ahead of the rest of the field. Reach for Perplexity when you want a sourced answer or want the web acted on, not because any single creative feature is best in class. ## Is Perplexity Free? Pricing Explained **Yes, Perplexity has a genuinely usable free tier, and the paid tiers mostly buy you higher limits and model choice.** For a lot of people the free plan is enough.
TierReported price (2026)What you get
Free$0Unlimited quick searches and citations, plus a small daily allotment of Pro searches. No model picker, no Labs.
Pro~$20 / month or ~$200 / yearUnlimited Pro searches, model choice, more Deep Research, larger file uploads, image generation
Max~$200 / monthThe highest limits, the newest features, and the most agent and Computer usage
Enterprise~$40 / seat / month and upTeam controls, internal document search, security and data-handling guarantees
A caveat: the exact free-plan caps move around, so the daily Pro-search number you read today may be different next month. Check the current limits in the app rather than trusting any single article. Run the free plan hard for a week. If you keep bumping into the daily Pro-search limit, the upgrade pays for itself. If you do not, stay free. One honest caveat from regular users: switching the underlying model on Pro changes the answer less than you would expect for most everyday questions, so do not upgrade only for the model picker. For a full plan-by-plan breakdown, see our [Perplexity pricing guide](https://geotoolbox.ai/blog/perplexity-pricing). ## Perplexity vs ChatGPT They overlap on a lot, so the real question is which one fits a given job, not which is "better." The cleanest way to think about it: Perplexity searches and cites, ChatGPT thinks and creates.
FactorPerplexityChatGPT
Default behaviorLive web search with citations on every answerWrites from training data; browses when asked
SourcesShown inline by defaultNot always shown
Best atResearch, fact-finding, current eventsWriting, brainstorming, coding, long-form
Weak spotLess suited to creative or extended back-and-forthCan sound confident without showing its sources
In practice many people run both: use Perplexity to gather sourced facts, then move to ChatGPT to draft and polish. If you already pay for ChatGPT, you do not strictly need Perplexity, but you give up the cited, current answers that are its whole point. We go deeper on the trade-offs in our [ChatGPT vs Perplexity](https://geotoolbox.ai/blog/chatgpt-vs-perplexity) comparison. And if you want to know whether Perplexity cites your own site, our [best Perplexity rank trackers](https://geotoolbox.ai/blog/best-perplexity-rank-tracker) guide compares the tools that measure it. ## Is Perplexity Accurate and Safe? The Honest Picture This is where a straight answer matters more than the marketing. Three things are worth knowing. **Accuracy: citations are a trust signal, not a guarantee.** Because every claim has a source attached, it is tempting to assume Perplexity is always right. It is not. Like any model, it can still [hallucinate](https://geotoolbox.ai/blog/ai-hallucinations), and the citation format can actually mask that. A known failure mode is a real link attached to a claim the source never actually made, which reads as verified but is not. The lawsuit from Dow Jones and the New York Post even alleged that Perplexity attributed made-up quotes to their articles. The takeaway is simple: the citations make verification easy, so use them. Click through on anything that matters, especially on long Deep Research reports. **Privacy: assume the consumer app trains on your data unless you turn that off.** On the free and Pro plans, your queries can be used to improve the models by default. There is an opt-out toggle in the account settings, and it generally applies going forward rather than retroactively. The practical rule is the one that applies to every consumer AI tool: do not paste anything confidential or regulated into it, and check the enterprise or API terms separately if you need stronger guarantees. Training policies shift often, so confirm the current setting in your account rather than trusting any single write-up. **Reputation: the scraping controversy is real, and the list of publishers suing keeps growing.** Perplexity has spent two years fighting over how it gathers content. The [New York Times sued it for copyright infringement](https://www.cnbc.com/2025/12/05/the-new-york-times-perplexity-copyright.html) in December 2025, escalating from an earlier cease-and-desist, and [CNN sued in May 2026](https://www.npr.org/2026/05/28/g-s1-124680/cnn-sues-ai-company-perplexity-alleging-it-violates-copyright-protections) over roughly 17,000 copied stories. Dow Jones, the New York Post, the BBC, several Japanese newspapers, Encyclopaedia Britannica, and [Reddit](https://www.cnbc.com/2025/10/23/reddit-user-data-battle-ai-industry-sues-perplexity-scraping-posts-openai-chatgpt-google-gemini-lawsuit.html) have all sued or threatened to. Most pointedly, in August 2025 Cloudflare reported that Perplexity used [undeclared "stealth" crawlers](https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/) that ignored robots.txt and impersonated a normal browser to get around blocks, and de-listed it as a verified bot. Perplexity has disputed the framing; The Verge keeps a running [account of the controversies](https://www.theverge.com/24187792/perplexity-ai-news-updates). In February 2026 the company dropped advertising and moved to a subscription-first model, saying it wanted to protect trust in the answer engine. None of this makes Perplexity unusable, but it is why some organizations keep it off their approved-tools list, and it is the backdrop to the part that matters most for site owners. ## What Perplexity Means for Your Brand's Visibility **When someone asks Perplexity about your category, it cites only a handful of sources, and if you are not one of them, you are invisible in that answer.** This is the shift that matters. The goal is no longer just to rank first on Google; it is to be one of the few sources the answer engine quotes. That discipline has a name, [answer engine optimization](https://geotoolbox.ai/blog/what-is-answer-engine-optimization), and it sits inside the broader practice of [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo). The good news is that Perplexity is unusually transparent about it. Because it shows its citations, you can ask it your own buyer's questions, read exactly which pages it pulled, and reverse-engineer the gap. That makes it the most measurable of the AI engines to optimize for, which is the whole premise of our guide to [getting cited in Perplexity](https://geotoolbox.ai/blog/perplexity-seo). Before any of that, there is a binary gate: reachability. Perplexity uses two crawlers, [PerplexityBot](https://geotoolbox.ai/glossary/perplexitybot) for its search index and [Perplexity-User](https://geotoolbox.ai/glossary/perplexity-user) for live fetches, and if a robots.txt or firewall rule blocks PerplexityBot, you drop out of the index it builds answers from. In our experience auditing brands across engines, this is the most common and most fixable reason a brand is missing from Perplexity answers, and it is invisible until you check. The citations themselves are the scoreboard, which is why we [track AI visibility per engine](https://geotoolbox.ai/blog/how-to-track-ai-visibility) rather than as one blended number. The open question for most businesses is the one you cannot answer by reading about Perplexity: what is it actually telling people about you, and where is it getting that from? geotoolbox tracks brand mentions and citations across the major AI engines, Perplexity included, and our free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) confirms in seconds whether PerplexityBot can even reach your pages. Start there, then [see what AI is citing about your brand](https://geotoolbox.ai/features/citation-interceptor) and close the gaps you find. ## Frequently Asked Questions ### What is the difference between Perplexity and ChatGPT? Perplexity is an answer engine built around live web search, so it cites its sources by default and is strongest for research and current facts. ChatGPT is built around generating text from its training, so it is stronger for writing, brainstorming, and coding. Many people use both: Perplexity to gather sourced facts, ChatGPT to draft. ### Is Perplexity AI free, and is Pro worth $20 a month? Yes, Perplexity has a free tier with unlimited basic searches and a small daily allotment of Pro searches. Pro costs about $20 a month and lifts those limits, adds model choice, and gives you more Deep Research. It is worth it if you regularly hit the free daily cap; if you do not, the free plan is fine. ### Does Perplexity have its own AI model, or does it use ChatGPT? Both. Perplexity builds its own Sonar models and also lets paid users pick frontier models from OpenAI, Anthropic, and Google to write the answer. It is an independent company, not owned by or built on OpenAI, so it is not "ChatGPT with sources." ### Can Perplexity replace Google Search? For research, comparison, and "explain this to me" questions, many people already use it instead of Google because it hands back a sourced answer rather than a list of links. For navigation, local results, maps, and shopping, Google still wins. In practice it complements Google more than it replaces it. ### Is Perplexity accurate, and can I trust its citations? It is accurate often enough to be useful, but it can still get things wrong, and a citation is not a guarantee. The cited source sometimes does not actually support the claim attached to it. Treat Perplexity as a fast first draft of the truth, and verify anything important before you rely on it. ### Who owns Perplexity, and can I buy its stock? Perplexity AI, Inc. is a private US company founded in 2022 by Aravind Srinivas and three co-founders, backed by investors including Nvidia and Jeff Bezos. Because it is private, there is no public stock to buy; any listing claiming to be "Perplexity stock" is not the company itself. ## Sources - Perplexity AI - Wikipedia (founders, funding, products, controversies) - `en.wikipedia.org/wiki/Perplexity_AI` - Perplexity received 780 million queries last month, CEO says - TechCrunch, June 2025 - `techcrunch.com/2025/06/05/perplexity-received-780-million-queries-last-month-ceo-says` - Airtel partners with Perplexity to power 360 million customers with Perplexity Pro - Airtel, July 2025 - `airtel.in/press-release/07-2025/airtel-partners-with-perplexity-powers-every-single-of-its-360mn-customers-with-perplexity-pro` - The New York Times sues Perplexity over copyright - CNBC, December 2025 - `cnbc.com/2025/12/05/the-new-york-times-perplexity-copyright.html` - CNN sues Perplexity over alleged AI copyright theft - NPR, May 2026 - `npr.org/2026/05/28/g-s1-124680/cnn-sues-ai-company-perplexity-alleging-it-violates-copyright-protections` - Reddit sues Perplexity over scraped user data - CNBC, October 2025 - `cnbc.com/2025/10/23/reddit-user-data-battle-ai-industry-sues-perplexity-scraping-posts-openai-chatgpt-google-gemini-lawsuit.html` - Perplexity is using stealth, undeclared crawlers to evade no-crawl directives - Cloudflare, August 2025 - `blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives` - Perplexity AI: the answer engine with a lot of question marks - The Verge - `theverge.com/24187792/perplexity-ai-news-updates` --- ## Chinese AI Models Compared: DeepSeek, Qwen, GLM, Kimi (2026) > An honest, current comparison of the top Chinese AI models - DeepSeek, Qwen, GLM, Kimi and more - covering capability, cost, safety, and AI visibility. - Canonical: https://geotoolbox.ai/blog/chinese-ai-models-compared - Published: 2026-06-27 · Updated: 2026-08-23 Chinese AI models went from a single viral release to a crowded, credible field in about eighteen months. If you are trying to make sense of DeepSeek, Qwen, GLM, Kimi, and the rest, the noise is the real obstacle: most coverage is either breathless or dismissive, and almost none of it is current, because the labs ship new versions every few weeks. This is a straight comparison of the major Chinese AI models as they stand in August 2026: who makes them, how good they actually are, whether they are safe to use, and what their rise means if you care about being found in AI answers. ## The Major Chinese AI Models at a Glance A year ago, "Chinese AI model" mostly meant DeepSeek. In August 2026 it means a crowded field of labs shipping frontier-grade models, most of them with downloadable [open weights](https://geotoolbox.ai/glossary/open-weights), and most of them far cheaper than the closed American systems. The US has since answered with its own frontier open model, Thinking Machines' [Inkling](https://geotoolbox.ai/blog/inkling-ai), but the field below is the one it is chasing. Here is the current lineup, with the version that is actually live as of this writing.
Model (Aug 2026)LabOpen weights?ContextBest atRough API price (input / 1M)
DeepSeek V4DeepSeek (Hangzhou)Yes~1MCheap reasoning and coding~$0.22 off-peak (Flash)
Qwen3.8-MaxAlibabaSmaller models yes; flagship weights published Aug 12, 2026 (text-only, bespoke license)1MMultilingual, agents, multimodal$2
GLM-5.2Zhipu / Z.ai (Beijing)Yes~1MCoding and agent work~$1.00
Kimi K2.6Moonshot AIYes~256KLong context, agent swarms~$0.80
MiniMax M3MiniMaxYes~1MLong context at low cost~$0.40
Ernie 5.1BaiduErnie 4.5 family yes~128KSearch-grounded, efficient~$0.55
Doubao Seed 2.1 ProByteDanceNo~256KHigh-volume production~$0.83
Prices move almost weekly and depend on tier, so treat them as orders of magnitude rather than quotes. The pattern that holds: the strongest Chinese models cost a fraction of GPT-5 or Claude Opus, and most of them you can download and run yourself. ## DeepSeek, Qwen, GLM, and Kimi: The Big Four Four labs do most of the heavy lifting in this conversation. They are the ones whose names show up in AI Overviews, the ones Western startups quietly build on, and the ones you are actually choosing between when you ask which Chinese model to use. ### DeepSeek (V4): The Cost Disruptor **DeepSeek is the lab that started the panic.** Its R1 reasoning model in January 2025 matched American frontier systems at a tiny fraction of the cost, became the most-downloaded free app in the US App Store within days, and helped wipe roughly a trillion dollars off US tech stocks in the late-January 2025 selloff. The lab sits in Hangzhou and is funded by the quantitative hedge fund High-Flyer, which is part of why it optimizes so aggressively for cost. Our [DeepSeek pricing guide](https://geotoolbox.ai/blog/deepseek-pricing) has the current V4 Flash and Pro rates in full. [DeepSeek's current flagship](https://geotoolbox.ai/blog/what-is-deepseek) is [V4](https://geotoolbox.ai/blog/deepseek-v4), which [TechCrunch reported was previewed in late April 2026](https://techcrunch.com/2026/04/24/deepseek-previews-new-ai-model-that-closes-the-gap-with-frontier-models/), built on the same mixture-of-experts efficiency the lab is known for; its flagship V4-Pro reached general availability on August 13, 2026. It is genuinely good at reasoning and agentic coding, and on price it is close to unbeatable: the Flash tier starts around twenty-two cents per million input tokens off-peak, where Western frontier models charge dollars. On August 16, 2026, DeepSeek replaced its old flat all-day rate with a peak/off-peak split that doubles the price during seven hours a day (peak hours 01:00-04:00 and 06:00-10:00 UTC), but even at peak it stays a fraction of Western pricing. If your question is "which Chinese model gets me the most capability per dollar," DeepSeek is usually the answer, and it is where most DeepSeek-versus-Qwen comparisons land on cost alone. ### Alibaba Qwen (Qwen3.8-Max): The Default Base Model If DeepSeek grabbed the headlines, [Qwen](https://geotoolbox.ai/blog/what-is-qwen) quietly became the foundation. Alibaba's Qwen family is the most-downloaded open model line in the world: [MIT Technology Review reports](https://www.technologyreview.com/2026/02/12/1132811/whats-next-for-chinese-open-source-ai/) it accounted for more than 30 percent of all Hugging Face model downloads in 2024, and by August 2025 Qwen derivatives made up over 40 percent of new language-model variants on Hugging Face, versus about 15 percent for Meta's Llama. Alibaba ships new Qwen models at a relentless pace, several a month. The practical takeaway is that when a developer anywhere fine-tunes "an open model," it is very often a Qwen underneath. The smaller and mid-size Qwen models ship with open weights under permissive licenses, which is what makes them the default starting point. The catch: Alibaba long treated its very largest models as a closed API rather than an open download, and only partly reversed that in August 2026. Qwen's strengths are breadth: strong multilingual coverage, agentic tool use, and a one-million-token context window. The current flagship, [Qwen3.8-Max](https://geotoolbox.ai/blog/qwen3-8-max), was released on August 3, 2026 at 2.4 trillion parameters with 95 billion active, priced at $2 per million input tokens and $6 per million output. Alibaba published its weights on August 12, 2026, the first open Max-tier Qwen, though as a text-only variant under a bespoke "Qwen3.8-Max" license rather than Apache 2.0. ### Zhipu GLM (GLM-5.2): The Frontier Challenger **[GLM-5.2](https://geotoolbox.ai/blog/what-is-glm-5-2) is the model that closed the gap.** It comes from Zhipu AI, a Beijing lab spun out of Tsinghua University that operates internationally as Z.ai, and it launched on June 13, 2026 under an MIT license with a one-million-token context window. On the independent BenchLM leaderboard it sits at the top of the Chinese field, and [Reuters reported](https://www.reuters.com/world/asia-pacific/after-anthropic-shutdown-chinas-zai-closes-frontier-gap-it-plans-dual-listing-2026-06-25/) it rivals the closed frontier models from OpenAI and Anthropic on coding and agent tasks at a fraction of the cost. Z.ai went public in Hong Kong in January 2026, and investors treated its decision to give the weights away as a strategic win rather than a loss. What sets GLM apart from DeepSeek is focus. Where DeepSeek optimizes for raw cost, GLM-5.2 is tuned for long-horizon software engineering: full-codebase debugging, autonomous self-correction, and the kind of multi-step agent runs that fall apart on weaker models. If your work is serious coding or agent orchestration and you want open weights, this is the current pick. ### Moonshot Kimi (K2.6): The Agent Specialist [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai), from Moonshot AI, is built for agents that run for hours. The K2 line ships with open weights under a modified MIT license, and the current K2.6 generation is engineered around long context and what Moonshot calls agent-swarm orchestration: coordinating many sub-agents across long background tasks rather than answering one prompt at a time. An earlier Kimi release landed at roughly one-seventh the price of Claude Opus and quickly became a heavily used model on agent platforms. In mid-June 2026 Moonshot followed K2.6 with Kimi K2.7-Code, a coding-specialized variant, and in July 2026 it launched [Kimi K3](https://geotoolbox.ai/blog/what-is-kimi-k3), a much larger 2.8-trillion-parameter model we cover in a separate explainer. K3 raised the ceiling on scale, but at roughly three-to-four times the price ($3/$15 per million in/out) and with open weights that followed shortly after launch (published July 27, 2026, under a bespoke Kimi K3 License), K2.6 remains the practical pick for most day-to-day agent work; our [Kimi K3 vs Claude](https://geotoolbox.ai/blog/kimi-k3-vs-claude) comparison weighs K3 against the Claude lineup task by task. On the consumer side Moonshot sells five subscription tiers metered in credits rather than messages, which we break down in our [Kimi pricing](https://geotoolbox.ai/blog/kimi-pricing) guide. In a head-to-head, comparing Kimi with DeepSeek, or Qwen with Kimi, usually comes down to job shape. DeepSeek is the cheapest generalist, Qwen is the most adaptable base, GLM is the coding frontier, and Kimi is the one you reach for when an agent needs to keep its footing across a very long, multi-step run. ## The Next Tier: MiniMax, Baidu Ernie, and ByteDance Doubao Three more labs round out the field and show up constantly in any serious comparison. **MiniMax** shipped M3 on June 1, 2026 with open weights. Its hook is architectural: a sparse-attention design that handles a one-million-token context window at far lower compute cost than a standard transformer, which lets it stay cheap while going long. It went public in Hong Kong the same week as Z.ai. **Baidu Ernie** is the veteran. Its current flagship, Ernie 5.1, launched in spring 2026 and reached the top of the Chinese field on the LMArena preference leaderboard while training on a fraction of the usual compute. The genuinely open piece is the older Ernie 4.5 family, which Baidu released under an Apache 2.0 license as a range of [mixture-of-experts](https://geotoolbox.ai/glossary/mixture-of-experts) models you can download and run. Ernie's edge is tight search grounding, which fits Baidu's search-engine roots. **ByteDance Doubao** is the odd one out: it is closed. The Doubao Seed 2.1 Pro line is served only through ByteDance's Volcano Engine cloud, and it is built for raw scale, with the family handling well over a hundred trillion token calls a day across ByteDance's apps. It is cheap and reliable for high-volume production, but you cannot download it, and the ByteDance name carries the same political baggage TikTok does in the US. One scope note before going further: this guide is about text and chat models, the ones you reason, write, and code with. The other place Chinese labs are winning is generative media, where they arguably lead the world. Kuaishou's Kling, MiniMax's Hailuo, ByteDance's Seedance, and Alibaba's Wan top most video-generation rankings, and Tencent's Hunyuan family rounds out the major labs not covered above. If your interest is AI video or image generation rather than chat, that is a different shortlist than the one here. ## Open Weights vs Open Source: What "Open" Really Means Almost every article on this topic calls these models "open source," and almost every one is being loose with the term. The distinction matters enough to get right, because it changes what you are actually allowed to do. Most Chinese models are open weights, not open source. The lab publishes the trained model file so you can download, run, and fine-tune it, but it does not publish the training data or the full training code. So you can use the model freely, but you cannot fully reproduce or audit how it was built. A genuinely [open-source AI](https://geotoolbox.ai/glossary/open-source-ai) model would release all three; an [open-weights](https://geotoolbox.ai/glossary/open-weights) model releases only the weights. It is the difference between being handed a working engine and being handed the engine plus the factory blueprints. For the model-by-model license breakdown, see our guide to [open weights vs open source](https://geotoolbox.ai/blog/open-weights-vs-open-source).
![Most Chinese frontier AI models ship open weights you can download and self-host. As of August 2026, DeepSeek V4, GLM-5.2, Kimi K3, MiniMax M3, and the smaller Qwen and Ernie models are open, and Alibaba opened Qwen3.8-Max's text weights on August 12 under a bespoke license; Baidu's Ernie 5.1, ByteDance's Doubao, Qwen3.7-Max, and the multimodal Qwen3.8-Max flagship stay API-only.](/blog/chinese-ai-models-compared/open-vs-closed.png)
As of August 2026, the open downloads span DeepSeek, GLM, Kimi K3, MiniMax, the smaller Qwen and Ernie models, and Qwen3.8-Max's text weights (bespoke license, opened August 12); Qwen3.7-Max, the multimodal Qwen3.8-Max flagship, Ernie 5.1, and Doubao stay closed.
"Open" also does not automatically mean "free for any use." The license attached to the weights decides that, and the licenses vary.
ModelLicenseCommercial useSelf-hostWhat is not released
DeepSeek V4MITYesYesTraining data and code
GLM-5.2MITYesYesTraining data and code
Kimi K3Kimi K3 LicenseYes, with conditionsYesTraining data and code
MiniMax M3MiniMax Community LicenseYes, with conditionsYesTraining data and code
Qwen (smaller models)Apache 2.0YesYesMultimodal flagship weights, training data
Ernie 4.5 familyApache 2.0YesYesErnie 5.x weights, training data
Doubao Seed 2.1ProprietaryAPI onlyNoThe weights themselves
The surprise for many people is that the permissive ones are genuinely permissive. MIT and Apache 2.0 are about as open as licenses get, with no user-count caps. That makes DeepSeek, GLM, and the open Qwen and Ernie models legally easier to build a commercial product on than [Meta's Llama](https://geotoolbox.ai/blog/what-is-meta-ai), whose license adds a large-user restriction. The thing to watch is not the openness of the small models but the trend at the top: Baidu's Ernie 5.1 and all of Doubao stay closed, and Alibaba opened only a text-only version of Qwen3.8-Max in August, keeping the multimodal flagship API-only. ## Are Chinese AI Models Safe? Privacy, Censorship, and Bans This is the question that actually stops people, and it deserves a straight answer. There are three real concerns, and for each one the answer depends heavily on how you use the model. The first is **data privacy**. When you use a hosted Chinese AI app or its native cloud API, your prompts are processed on servers in China, and China's 2017 National Intelligence Law can compel companies to hand data to the state. That is a genuine problem for regulated work, and it is why many employers already restrict the consumer DeepSeek and Qwen apps. The key distinction is hosting: that risk attaches to the hosted service, not to the open weights. Download the weights and [run them on your own hardware](https://geotoolbox.ai/blog/run-llm-locally) or a Western cloud, and your data never leaves your environment. Same model, different data path. The second is **censorship**. Hosted Chinese assistants refuse or deflect on politically sensitive topics like Tiananmen, Taiwan, and the Uyghurs, and a September 2025 report from the US government's AI safety body, [covered by NBC News](https://www.nbcnews.com/tech/innovation/silicon-valley-building-free-chinese-ai-rcna242430), found DeepSeek models showed weakened safety protocols and a measurable lean toward pro-Chinese framing compared with US models. For most marketing or coding work this never comes up, but if you are generating content where political neutrality matters, test for it. Self-hosting an open-weight version reduces but does not fully remove the bias baked in during training. The third is **bans**. US federal agencies and several states restrict DeepSeek on government devices, and a few allied countries have moved against the hosted apps. These bans almost always target the hosted service on official devices, not the underlying open weights running in a private deployment. So "is DeepSeek banned" depends entirely on who and where you are. Restriction is not a one-way street, either. In June 2026 the US itself briefly made Anthropic [pull its top Claude models offline](https://geotoolbox.ai/blog/fable-5-ban) (Fable 5 and Mythos 5) to comply with an order restricting foreign access, a restriction lifted on July 1, 2026, but one that, while it lasted, pushed developers in other countries toward exactly these Chinese open-weight models. [CrowdStrike has reported](https://www.crowdstrike.com/en-us/blog/crowdstrike-researchers-identify-hidden-vulnerabilities-ai-coded-software/) that DeepSeek-R1 emits up to 50 percent more insecure code when prompts include politically sensitive triggers, which most analysts read as a training side effect rather than deliberate sabotage. Review the output, and keep these models away from government or safety-critical systems, the same as you would any model you cannot fully audit. ## How They Compare to ChatGPT, Claude, and Gemini Chinese models have closed most of the gap on most tasks, while a real gap remains at the very top. On public benchmarks for coding, math, and reasoning, the best Chinese models now trade blows with GPT-5, Claude, and Gemini. On cost the gap is wide: the Chinese options routinely run a fraction of the price, DeepSeek's twenty-two cents against dollars for the Western flagships, which is the single biggest reason Western developers reach for them. Two cautions matter. First, a lot of "beats GPT-5" framing comes from the labs themselves, and independent observers still credit US labs with a lead at the frontier, describing the Chinese players as fast, well-funded fast-followers, even as they have drawn level on specific axes like coding and agent benchmarks. Second, benchmarks are easy to over-read. Many leaderboards have thin English-language coverage, and there are ongoing questions about test contamination, so a single eye-catching score is weaker evidence than a consistent pattern across independent evaluations. It is also worth remembering that these labs still train largely on Nvidia hardware and US cloud infrastructure, so "independent Chinese AI" is a simplification. For practical work, the split is clean. If you want maximum capability and have the budget, the top closed American models still have an edge on the hardest tasks and the smoothest tooling. If you want most of that capability at a tenth of the cost, or you need to run a model in your own environment, the Chinese open models are now a serious answer rather than a curiosity. If you are weighing one against a specific Western system, our [Gemini vs ChatGPT comparison](https://geotoolbox.ai/blog/gemini-vs-chatgpt) lays out the same framework for the closed side. ## What the Rise of Chinese AI Means for Your Visibility Here is the part the benchmark roundups skip, and it is the one that matters if your job is marketing rather than machine learning. Chinese open models are no longer a niche. They grew from essentially zero to around 30 percent of usage on OpenRouter, a major model-routing service, in little over a year, and roughly 80 percent of startups building on open-source stacks now run on them. [MIT Technology Review's framing](https://www.technologyreview.com/2026/02/12/1132811/whats-next-for-chinese-open-source-ai/) is the right one: these models have become infrastructure for global AI builders. Microsoft, for one, [added DeepSeek's R1 to its Azure AI Foundry catalog](https://azure.microsoft.com/en-us/blog/deepseek-r1-is-now-available-on-azure-ai-foundry-and-github/) within days of the model's release. That changes the visibility picture in two ways. The first is that a growing share of the agents and applications reading your site are powered by these models, so [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) increasingly runs through Chinese model weights even on Western products. The second is that the big Chinese assistants, DeepSeek, Kimi, Doubao, Ernie, and Qwen's Tongyi, are their own answer engines with hundreds of millions of users, and almost nobody in Western marketing is thinking about whether they appear there. In our experience at geotoolbox, the brands that stay visible as the model field splinters are the ones that treat AI reachability and citation as a first-class channel rather than an afterthought. Our [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) covers 34 crawler user-agents, including ByteDance's, so you can at least see which AI systems your robots.txt currently lets in. The open question nobody has answered well yet, ourselves included, is how consistently these Chinese engines cite and link back to the Western sources they draw on, which is exactly the kind of gap worth watching rather than guessing about. ## Which Chinese AI Model Should You Use? There is no single winner, because the right pick depends on the job. Here is the short version. For the global field including Llama, Gemma, and gpt-oss, our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking covers all ten.
If you want to...PickWhy
Pay as little as possible per tokenDeepSeek V4Frontier-level output at the lowest price
Do serious coding and agent work, openGLM-5.2Tops the Chinese field on coding and long agent runs
Run long, multi-step agent jobsKimi K2.6Built for long context and agent-swarm orchestration
Fine-tune your own modelQwen (open sizes)The most-adopted, best-supported open base
Run locally on modest hardwareA small Qwen or DeepSeekSmaller open variants fit consumer GPUs
Avoid China-hosted data entirelyAny open model, self-hostedOpen weights on your infrastructure keep data local
A practical note on access: you rarely need a Chinese account to use these. The open models are on Hugging Face, and Western-friendly gateways like OpenRouter let you call most of them through one API without touching a native cloud. That avoids a Chinese account and cloud while keeping the price advantage, though a gateway still routes your prompt through a third party, so check its retention terms if data residency matters. So "which Chinese AI model is best" is the wrong question. The field is good enough now that the better question is which one fits your specific job, budget, and risk tolerance, and the table above is where to start. ## The Models Will Keep Changing. Your Visibility Strategy Shouldn't Have To The specific version numbers in this guide will be stale within months, because that is the pace these labs ship at. What will not change as fast is the underlying shift: capable AI is getting cheaper, more open, and more fragmented across providers, and a rising share of it now runs on Chinese weights. For anyone whose livelihood depends on being found, that fragmentation is the real story. When a dozen engines, Western and Chinese, can all answer a buyer's question, the brands that win are the ones cited across them, not the ones optimized for a single search box. That is the problem geotoolbox is built for: our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) tracks where you appear, and where you are missing, across ChatGPT, Perplexity, Gemini, Claude, Google AI Overviews, Bing Copilot, and Grok, so you can fix the gaps instead of guessing at them. Start with the engines that already send you buyers, and keep one eye on the Chinese ones coming up fast behind them. ## Frequently Asked Questions ### What is the best Chinese AI model right now? There is no single best one. On independent coding and agent leaderboards GLM-5.2 leads the Chinese field, DeepSeek V4 wins on price, Qwen is the most-adopted open base, and Kimi K2.6 is strongest for long agent runs. The right pick depends on whether you care most about cost, coding, fine-tuning, or agent work. ### Is DeepSeek banned in the US? Not for the general public. US federal agencies and several states ban the DeepSeek app on government devices over data-security concerns, and a few other countries have restricted the hosted app. Those bans target the China-hosted service, not the open-weight model, which anyone can still download and run privately. ### Are Chinese AI models safe to use, and does my data go to China? If you use the hosted apps or native cloud APIs, your prompts are processed in China and can be subject to Chinese data laws, which is a real concern for sensitive work. If you self-host the open weights on your own hardware or a Western cloud, your data never reaches China. A gateway like OpenRouter avoids a Chinese account and cloud, though your prompt still passes through a third party, so check its routing and retention terms. ### What is the difference between open-weight and open-source? Open-weight means the lab publishes the trained model so you can download, run, and fine-tune it. Open-source would also include the training data and code needed to rebuild it from scratch. Most of the leading Chinese models, like DeepSeek and GLM, are open-weight but not fully open-source, so you can use them freely without being able to fully audit how they were made. ### Can I run Chinese AI models locally or offline? Yes, for the open-weight ones. Smaller DeepSeek, Qwen, and GLM variants, around 7 to 14 billion parameters and quantized, run on a single 8 to 16 GB consumer GPU through tools like Ollama. The trillion-parameter flagships need multi-GPU or datacenter hardware, though even that ceiling is starting to move: a July 2026 proof-of-concept called [Colibri](https://www.tomshardware.com/tech-industry/artificial-intelligence/colibri-proof-of-concept-gains-frontier-level-1-5-tb-ai-model-novel-approach-runs-on-only-25gb-of-ram-and-shows-promise-for-local-ai-setups) ran GLM-5.2, natively a 1.5 TB model, on a 25 GB consumer machine by streaming its experts from disk, but at roughly 0.05 to 0.1 tokens per second (one token every ten to twenty seconds) it is a glimpse of where local inference is heading, not a daily driver yet. Running locally is the cleanest way to remove the hosted data-residency risk, though training-time bias or refusal behavior can still surface. ### Which Chinese AI models do ChatGPT, Gemini, and Perplexity cite? That is a different question from which models are good. Western answer engines cite sources based on their own retrieval, not on which model a lab built, so being mentioned by ChatGPT or Gemini is about your content earning the citation. The newer wrinkle is that Chinese assistants like DeepSeek and Kimi are their own answer engines, and how they cite Western sources is still an open and largely unmeasured question. ## Sources - DeepSeek previews new AI model that closes the gap with frontier models - TechCrunch - `techcrunch.com/2026/04/24/deepseek-previews-new-ai-model-that-closes-the-gap-with-frontier-models` - What's next for Chinese open-source AI - MIT Technology Review - `technologyreview.com/2026/02/12/1132811/whats-next-for-chinese-open-source-ai` - More of Silicon Valley is building on free Chinese AI - NBC News - `nbcnews.com/tech/innovation/silicon-valley-building-free-chinese-ai-rcna242430` - After Anthropic shutdown, China's Z.ai closes frontier gap - Reuters - `reuters.com/world/asia-pacific/after-anthropic-shutdown-chinas-zai-closes-frontier-gap-it-plans-dual-listing-2026-06-25` - Hidden vulnerabilities in AI-coded software - CrowdStrike - `crowdstrike.com/en-us/blog/crowdstrike-researchers-identify-hidden-vulnerabilities-ai-coded-software` - DeepSeek R1 is now available on Azure AI Foundry - Microsoft Azure - `azure.microsoft.com/en-us/blog/deepseek-r1-is-now-available-on-azure-ai-foundry-and-github` - Artificial intelligence industry in China - Wikipedia - `en.wikipedia.org/wiki/Artificial_intelligence_industry_in_China` - DeepSeek API documentation - DeepSeek - `api-docs.deepseek.com` - Open models on Hugging Face - Hugging Face - `huggingface.co/models` - Colibri proof-of-concept runs GLM-5.2's 1.5-TB model on 25GB of RAM - Tom's Hardware, July 2026 - `tomshardware.com/tech-industry/artificial-intelligence/colibri-proof-of-concept-gains-frontier-level-1-5-tb-ai-model-novel-approach-runs-on-only-25gb-of-ram-and-shows-promise-for-local-ai-setups` --- ## Gemini SEO: How to Get Cited in Google Gemini > Gemini SEO means getting cited in Google Gemini's answers. How grounding, the Google index, Google-Extended, and the four Gemini surfaces actually work. - Canonical: https://geotoolbox.ai/blog/gemini-seo - Published: 2026-06-27 · Updated: 2026-07-20 Gemini SEO is not a new playbook so much as your existing Google SEO aimed at a new target: getting your brand named when someone asks Google Gemini a question, instead of just ranking a blue link. Most guides skip the reason it works. Gemini does not run its own index of the web; it grounds its answers on Google's Search index, the same one behind your rankings, and it shows up in four different places, from the [Gemini app](https://geotoolbox.ai/blog/what-is-gemini) to the AI answers inside Search. Get a few mechanics right and the same body of work feeds all four. This guide is the specifics, current as of July 2026. ## What "Gemini SEO" Actually Means Search "Gemini SEO" and most of what you find is about using Gemini to do your SEO work: drafting content, building schema, auditing a site. This guide is the other job. It is about getting your brand cited when someone asks Google Gemini a question, and that is now a distinct surface worth optimizing for because Gemini shows up in far more places than the chat app. "Gemini" is not one product. It is a family of models that powers several different answer surfaces, and each one can name your brand or skip it: - The **Gemini app** at gemini.google.com, where people chat directly - **AI Overviews**, the synthesized answer boxes at the top of Google Search - **[AI Mode](https://geotoolbox.ai/blog/google-ai-mode-seo)**, Google's conversational search tab that fans a question into many parallel sub-queries - **Gemini grounding** in the API, where other companies build apps on Google's search-grounded model What ties all four together is where they get their facts, and that is where every tactic below begins. It is also why optimizing for Gemini is mostly your existing [Google SEO and GEO](https://geotoolbox.ai/blog/what-is-geo) work, aimed at four AI surfaces instead of ten ranked links. ## Where Gemini Gets Its Answers When Gemini answers a question that needs facts, it does not invent them from memory. It runs a search, reads the results, and writes an answer that cites them. This is [grounding](https://geotoolbox.ai/glossary/grounding): the model is anchored to live, retrieved web content instead of relying only on what it learned in training. Google describes the developer version plainly in its [Gemini API grounding docs](https://ai.google.dev/gemini-api/docs/google-search), which connect the model to real-time Google Search and return inline citations to the pages it used. The retrieval pool is the important part. Gemini grounds on Google's own Search index, the corpus Googlebot already built. There is no separate "Gemini index" of the web that you submit to or optimize for independently. If your page is indexed, it is eligible for those grounded answers. If it is not indexed, it is invisible to every Gemini surface, no matter how good the content is. This is why the work overlaps so heavily with classic [AI search optimization](https://geotoolbox.ai/blog/how-does-ai-search-work). One twist changes how you structure a page. Gemini, and especially AI Mode, uses [query fan-out](https://geotoolbox.ai/blog/query-fan-out): it breaks one question into several related sub-questions, runs them in parallel, and assembles the answer from the best passage for each. So a single page that cleanly answers the main question plus the obvious follow-ups can be cited several times in one response, while a page that answers only the headline question gets cited once or not at all. You are not optimizing for one query anymore. You are optimizing for a small tree of them. ## The Crawler Question: Googlebot, Google-Extended, and the Trap This is where most "Gemini SEO" advice goes wrong, and the mistake can quietly remove you from the answers you want. There is no dedicated "GeminiBot" search crawler to allow. Discovery runs on **Googlebot**, the same crawler that indexes you for normal Search. The bot people confuse it with, [Google-Extended](https://geotoolbox.ai/glossary/google-extended), is not a search crawler at all. It is a control token in robots.txt that governs whether your content trains Google's generative models and grounds the Gemini app and Vertex AI.
TokenWhat it controlsBlock it and...
GooglebotCrawling and indexing for Google Search: the pool AI Overviews and AI Mode draw from, and the prerequisite index for the Gemini app and API tooYou fall out of the index, so you cannot be cited in AI Overviews, AI Mode, or grounded Gemini answers
Google-ExtendedWhether your content trains Google's models and grounds the Gemini app and the API (Vertex AI)You opt out of training and of grounding in the Gemini app and the API, with no effect on your Search ranking or on AI Overviews and AI Mode eligibility
"GeminiBot"Does not exist as a separate search crawlerNothing to block; do not add invented user-agents to robots.txt
Google is explicit that these are separate doors. Its documentation on [AI features and your website](https://developers.google.com/search/docs/appearance/ai-features) states that AI Overviews and AI Mode are built into Search, so "robots.txt directives for Googlebot is the control for site owners to manage access," and that you limit what they show with the standard `nosnippet`, `data-nosnippet`, `max-snippet`, or `noindex` controls. Google-Extended is for "some of Google's other systems," not Search. Google's own [crawler documentation](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers) puts the load-bearing line plainly: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." So training and citation are two switches, not one. Plenty of brands block Google-Extended to opt out of AI training, a perfectly reasonable choice, and assume they have "blocked Gemini." Not quite. Blocking Google-Extended does drop you from the two surfaces it gates, the Gemini app and the API, but you stay fully eligible for AI Overviews and AI Mode, because those ride Googlebot. The damaging mistake is the reverse: a broad firewall or robots rule that blocks Googlebot, which pulls you out of the index every Gemini surface draws from. Our [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) fetches your robots.txt server-side and shows which of 34 AI crawlers, Google-Extended and Googlebot included, you are allowing or blocking, so you can confirm you have not closed the wrong door. ## The Four Gemini Surfaces, and How Each Picks Sources Because all four surfaces ground on the same Google index, the foundation is shared, but they differ in how they retrieve and where you check whether you won. Treat them as one optimization target with four scoreboards.
![The four Google Gemini surfaces, all grounded on one Google Search index.](/blog/gemini-seo/four-gemini-surfaces-one-index.png)
Googlebot builds the index; all four Gemini surfaces draw their citations from it.
SurfaceWhat powers itHow you influence itWhere to check
Gemini appThe Gemini model, grounded on Google Search when a question needs fresh factsBe indexed and be the clean, current answer; off-site mentions help corroborate the brand tooAsk Gemini your questions in a signed-out session and read the cited links
AI OverviewsThe standard Search index, summarized at the top of resultsBe indexed and extractable; position is a weak signal, so see the dedicated playbookSearch Console, bundled under "Web"; manual SERP checks
AI ModeThe same index, with heavier query fan-out and deeper synthesisCover the question neighborhood on one page so you win multiple sub-queriesManual checks; Search Console does not break it out separately yet
Gemini API groundingDeveloper apps calling Gemini with Grounding with Google Search enabledBe indexed and keep Google-Extended unblocked; this surface is gated by Google-Extended, like the appInspect the inline url_citation annotations the API returns
The first two are familiar territory. The [Gemini app](https://geotoolbox.ai/blog/what-is-gemini) behaves like the other assistant chatbots: it searches, it cites, and the citation is a link the user can click. AI Overviews is a Search feature with its own established craft, which we cover in depth in [how to get cited in Google AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo). [AI Mode](https://geotoolbox.ai/glossary/google-ai-mode) is the one to watch. It leans hardest on fan-out, running many sub-queries behind a single conversational thread, which means the page that answers a whole topic cleanly tends to be pulled in repeatedly. The fourth surface is the one almost no guide mentions: the **Gemini API**. When another company builds a feature on Gemini with grounding turned on, Google's [API documentation](https://ai.google.dev/gemini-api/docs/google-search) shows it returns inline `url_citation` annotations pointing back to source pages. You do not optimize for it separately. You appear there for the same reason you appear in the app: you are in the Google index, you have not blocked Google-Extended, and you are the best passage for the sub-question. Same work, one more place it pays off. ## First, Can Gemini Reach Your Pages? Every tactic below assumes one thing: your page is in Google's index and the content Gemini needs is actually in the HTML it reads. That assumption fails more often than people expect, and when it fails, nothing else matters. Start with indexation. If a page is not indexed in Google Search, it cannot surface in any Gemini surface, full stop. Check coverage in Search Console before you spend an afternoon on formatting. A page that returns a soft 404, sits behind a `noindex` you forgot about, or never got crawled is not a citation candidate. Then check rendering. Gemini's grounding works from what Google has indexed, and a search-time read favors content that exists in the raw HTML. If your key facts, prices, or definitions are injected client-side by JavaScript, a crawler or grounding fetch can land on a near-empty shell and move on to a competitor whose answer sits in the served markup. Server-rendered, text-first pages are easier to extract than app-shell pages that paint the important parts after load. This is the reachability gate that governs [optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) on every engine, not just Gemini. Confirm the page is indexed, confirm the answer is in the HTML, and confirm no robots or firewall rule is blocking Googlebot. Only then is it worth tuning the content itself. ## What Gets You Cited in Gemini With the page reachable, citation comes down to being the clearest, most credible source for a specific claim. The fundamentals match the rest of AI search, with a few Gemini-shaped emphases. 1. **Lead with the answer.** State the fact directly in a self-contained sentence near the top of the relevant section. Gemini lifts passages, so a clean claim it can quote in one line beats the same point spread across three paragraphs. Aggarwal et al.'s [generative engine optimization study](https://arxiv.org/abs/2311.09735) found that adding citations, quotations, and statistics improved a source's visibility in their benchmark by up to 40 percent. This is the core of [writing pages LLMs cite](https://geotoolbox.ai/blog/ai-content-optimization). 2. **Cover the neighborhood.** Because of query fan-out, answer the obvious follow-up questions on the same page, each under its own clear heading. The page that satisfies the whole question tree gets cited across the thread; the page that answers only the headline gets cited once. 3. **Be a recognized entity.** Gemini leans on Google's [Knowledge Graph](https://geotoolbox.ai/glossary/knowledge-graph) and entity systems to know who you are. Consistent naming across the web, an Organization profile, and `sameAs` links to authoritative profiles help Google treat you as one verified entity rather than a stranger. This is where [entity SEO](https://geotoolbox.ai/blog/entity-seo) pays off. 4. **Make it verifiable and current.** Back claims with specific numbers, named sources, and visible dates. On changing topics, a fresh page is a stronger candidate than stale content, and dated facts help on exactly the queries that trigger a live search. There is no guaranteed pickup time; grounded retrieval favors recently crawled pages, so a visible last-updated date and a real refresh earn their keep. 5. **Earn off-site consensus.** Gemini synthesizes from many sources, so being named accurately in roundups, reputable reviews, and community threads matters as much as your own page. Video is its own lane here: Gemini is multimodal and YouTube is Google-owned, so a relevant video can be pulled into an answer the way a page is. AI engines cite what multiple sources agree on, not what a brand claims about itself. This is the [E-E-A-T signal](https://geotoolbox.ai/blog/eeat-ai-search) that moves you from eligible to cited. 6. **Structure for extraction.** Clear headings, short lists, comparison tables, and a tight definition near the top all make a passage easier to lift cleanly. One pattern surprises SEOs. Ranking number one does not guarantee the citation. In the scans we run at geotoolbox, AI Overviews and AI Mode regularly cite a clean, well-structured page ranking well down the results over the top-ranked result that buries its answer under a long intro. Position helps you qualify; extractable structure wins the slot. If you rank well but never get cited, the cause is usually not authority. Your best fact is just too buried to lift. ## The "Gemini SEO" Myths to Skip A few tactics get sold as Gemini requirements that Google's own documentation contradicts. Skip them and spend the time on the list above. **You need an llms.txt file.** You do not. Google states directly that you "don't need to create new machine readable files, AI text files, or markup" to appear in these features. There is no evidence Gemini reads llms.txt, so treat it as unproven, not a requirement. **Schema markup triggers citations.** It does not, at least not directly. The same Google guidance is explicit that there is "no special schema.org structured data that you need to add" to appear in AI Overviews or AI Mode. Standard schema still earns its place by helping Google understand your entities and by qualifying you for rich results, but it is plumbing, not a citation button. Do not expect a FAQ block to pull you into an answer on its own. **You have to be in the training data.** Also false, and this is the one the crawler section already debunked. Gemini grounds on the live index for current questions, so a page published yesterday can be cited today as long as it is crawled and indexed. Being part of a past training run is neither necessary nor sufficient. ## How to Measure Gemini Visibility Be honest with yourself about the measurement gap before you build a dashboard, because Google's tools do not make this easy yet. Search Console folds AI Overviews into the regular "Web" search type and does not break out AI Mode separately, so you cannot cleanly isolate how much of your impression count comes from an AI answer. Referral data is worse: clicks from Gemini answers often arrive stripped of a clear referrer, so a meaningful share of Gemini-driven sessions land in your analytics as direct traffic rather than anything labeled "gemini." To recover the clicks that do carry a referrer, set up a GA4 custom channel that matches the gemini.google.com referrer, and treat it as a floor, since it misses the zero-click majority. So measurement is mostly manual, and it works like this: 1. **Build a prompt set.** Write down the 10 to 20 real questions your customers would ask Gemini in your category, the ones an answer about you should appear in. 2. **Run them and read the citations.** Ask each in a signed-out Gemini session, in AI Mode, and watch the AI Overview on the same query in Search. Record whether you are cited and, just as useful, who is cited instead. 3. **Sample, do not snapshot.** Gemini is non-deterministic; the same prompt can cite different sources on different runs. Ask each question a few times before you conclude you have won or lost a slot, and track your citation rate rather than a single yes or no. 4. **Re-run on a schedule.** Monthly checks turn into a trend you can act on, and the list of answers you are missing from becomes a ranked to-do list of passages to improve. Doing that by hand across surfaces and runs gets old fast, which is the gap our tools are built for. geotoolbox's [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) tracks brand citations across Gemini and Google AI Overviews alongside the other major engines, and the [Domain Overview](https://geotoolbox.ai/features/domain-overview) turns repeated checks into a visibility baseline over time. Whether you automate it or keep a spreadsheet, the discipline is the same: sample real prompts, log who gets cited, and close the gaps one passage at a time. For the broader workflow, see our guide on [measuring AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). ## Is Gemini Worth Optimizing For? More than almost any other AI engine, and not because the Gemini app is the biggest chatbot. The reach comes through Search. At Google I/O in May 2026, Sundar Pichai said [AI Overviews had passed 2.5 billion monthly users](https://www.cnbc.com/video/2026/05/19/google-ceo-pichai-ai-overviews-now-has-over-2-point-5-billion-monthly-users.html), and AI Mode crossed a billion monthly users in its first year. Those are Gemini-powered answers sitting on top of the search results your audience already runs. Being cited there is visibility at a scale no standalone assistant matches today. It also reframes what a win looks like. As AI answers absorb more of the clicks, the payoff shifts from a visit to a mention: being named in the answer, plus the branded searches and direct visits that follow when someone trusts what they read. Plan for presence in the answer, not only traffic to the page. The better argument is the marginal cost. The four surfaces share the same Google foundation, so optimizing for Gemini is not a separate project. It is the SEO and [GEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) work you already do: a reachable, indexed site, clear and current answers, real entity authority, and off-site consensus, with Google-Extended left open if you want the app and API surfaces too. So stop treating "Gemini SEO" as its own discipline. It is the highest-reach payoff of doing AI search optimization properly, the same way [getting cited in Claude](https://geotoolbox.ai/blog/claude-seo), [Perplexity](https://geotoolbox.ai/blog/perplexity-seo), or [Microsoft Copilot](https://geotoolbox.ai/blog/copilot-seo) is. ## Frequently Asked Questions ### Does Gemini cite sources? Yes, when it grounds a question on a live search. Gemini answers built from retrieval carry inline citations to the pages they used, and the Gemini API returns the same as inline citation annotations with the source URLs. When it answers from memory alone on an evergreen question, there are no per-claim sources, which is why getting cited means earning a place in the grounded answers. ### Is there a Gemini crawler I need to allow? No dedicated one. Discovery runs on Googlebot, the same crawler that indexes you for normal Search, so if Googlebot can reach you, Gemini's surfaces can cite you. Google-Extended is a training and grounding control, not a search crawler, and there is no "GeminiBot" user-agent to add to robots.txt. ### Does blocking Google-Extended hurt my Google rankings or AI Overviews? No. Google has confirmed Google-Extended is not a ranking signal and does not affect your inclusion in Search. Blocking it opts you out of model training and of grounding in the Gemini app and the API, while leaving you fully eligible for AI Overviews and AI Mode, which run on Googlebot. ### How is Gemini SEO different from Google SEO? It shares the same foundation, since Gemini grounds on the Google index, but the goal shifts from position to citation. That raises the weight on an answer-first, extractable structure, on covering the follow-up questions for query fan-out, and on off-site consensus, while lowering the weight on simply ranking first. ### How do I check if Gemini mentions my brand? Ask Gemini your category questions in a signed-out session, repeat across AI Mode and the AI Overview on the same query, and record whether you are cited and who is cited instead. Because answers vary between runs, sample each prompt a few times and track your citation rate rather than a single result. ### Can I see Gemini referral traffic in analytics? Partly. Clicks from the Gemini app can arrive with a gemini.google.com referrer, but most AI answers are read without a click, and many sessions lose the referrer and show up as direct traffic. Treat referral data as a floor, not a full count, and lean on manual citation checks for the real picture. ## Start with Reachability That is the whole job: make your pages reachable, then make them the cleanest answer to the question. Do those two things and Gemini follows the rest of your AI visibility work, no separate playbook required. Start at the gate the citation guides skip: can Googlebot actually reach and render your pages? geotoolbox's free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows whether your robots.txt allows Googlebot and Google-Extended, the free [AI-Readiness Score](https://geotoolbox.ai/tools/ai-readiness) checks the wider crawler-access foundations, and the [Content Analyzer](https://geotoolbox.ai/features/content-analyzer) grades how citable a page is once they are in place. Confirm the door is open, make your best pages the clearest current answer to the questions you want, and you give every Gemini surface its best reason to pick you. ## Sources - AI features and your website - Google Search Central - `developers.google.com/search/docs/appearance/ai-features` - Grounding with Google Search - Google AI for Developers (Gemini API) - `ai.google.dev/gemini-api/docs/google-search` - Google-Extended and the Google crawlers - Google Search Central - `developers.google.com/search/docs/crawling-indexing/google-common-crawlers` - GEO: Generative Engine Optimization - Aggarwal et al., KDD 2024 - `arxiv.org/abs/2311.09735` - Google CEO Pichai: AI Overviews now has over 2.5 billion monthly users - CNBC, Google I/O 2026 - `cnbc.com/video/2026/05/19/google-ceo-pichai-ai-overviews-now-has-over-2-point-5-billion-monthly-users.html` --- ## Grok Pricing in 2026: Plans, Free Tier, and Is It Worth It? > Grok pricing explained, current to August 2026: the free tier, SuperGrok ($10/$30), Heavy, X Premium vs standalone, API token costs, and how to use Grok. - Canonical: https://geotoolbox.ai/blog/grok-pricing - Published: 2026-06-27 · Updated: 2026-08-13 Grok is free to start, and for a lot of people that is the end of the story. If you want more, the standalone paid plans run $10 a month for SuperGrok Lite, $30 for SuperGrok, and $300 for SuperGrok Heavy, though Heavy is discounted to $99 a month for three months as of this writing. There are also team plans at $30 per seat, and a developer API billed per token from about $1 per million. The trouble is that Grok is sold two overlapping ways, the prices moved more than once in 2026, and most Grok pricing guides you will find still quote tiers that have changed. Below is every current Grok price, reconciled from each vendor's own pricing page, plus the question the numbers exist to answer: which plan, if any, you should actually pay for. One quick disambiguation first, because the search results mix them: this is about xAI's Grok chatbot, not the unrelated "Grok" crypto token. ## How Much Does Grok Cost? Every Plan at a Glance Here is the whole lineup in one place, at US prices as of July 2026.
PlanPrice (US, per month)What it isBest for
Free$0Grok on the web and in the X app, usage-cappedCasual questions and light use
SuperGrok Lite$10Standalone Grok with higher limits, lighter featuresRegular users on a budget
SuperGrok$30The main paid tier; 7-day free trialDaily and power users
SuperGrok Heavy$99 for 3 months, then $300Maximum limits, multi-agent modeHeavy and professional users
Grok Business$30 per seatTeam management, no training on your dataTeams and companies
Grok EnterpriseCustomSSO, SCIM, custom data retentionLarge organizations
Two things cause most of the confusion, so it is worth getting them straight before you pay for anything. Grok is sold two completely different ways for individuals. One is a standalone subscription you buy directly at [Grok's plans page](https://grok.com/plans), which is the SuperGrok line in the table above. The other is bundled into an X Premium subscription, where Grok comes attached to the social platform. They are not the same product, and people routinely overpay for one when they wanted the other. We untangle that below. SuperGrok also offers annual billing at a discount, around $300 a year, via the yearly toggle on the plans page. The per-tier annual figures move, so confirm them there before you prepay. The second point: the developer API is separate again. If you are wiring Grok into your own software, you do not buy a plan at all, you pay per token. Most people reading a pricing page want the consumer tiers above; if you are a developer, skip down to the API section. For the bigger picture of what the product actually is, our guide to [what Grok is](https://geotoolbox.ai/blog/what-is-grok) covers the models and history in plain terms.
![Grok's official plans page showing SuperGrok Lite at $10, SuperGrok with a 7-day free trial then $30, and SuperGrok Heavy at $99 for three months then $300, in June 2026.](/blog/grok-pricing/grok-plans-june-2026.png)
Grok's standalone plans on grok.com, June 2026, including the live SuperGrok Heavy promo.
## Is Grok Free? What the Free Tier Actually Gives You Yes. Grok has a free tier, and you do not need an X account to use it. Sign in at grok.com with a Google account or email and you can start chatting and getting real-time answers pulled from the web and X. The catch is the usage cap. Free accounts get a limited number of messages in a rolling window before Grok asks you to wait or upgrade. xAI does not publish an exact figure and it shifts with demand, but the pattern users report is roughly ten prompts every couple of hours, with the most capable model reserved for paying tiers. The window is rolling, not a daily reset, so if you hit the wall you wait it out rather than losing access for the day. Ask a few questions a day and you will rarely notice the cap. Lean on Grok for real work and you will hit it fast, which is the signal to consider paying. The free tier has also been shrinking. The clearest example is the [Grok Imagine](https://geotoolbox.ai/blog/grok-imagine) creator: its image and video generation was effectively free for casual users at launch, then xAI [restricted it to paying subscribers](https://www.calcalistech.com/ctechnews/article/rkbynj99bx) in early 2026, a change it framed as temporary. Image and video creation through Grok Imagine now sits behind the paid tiers, starting with SuperGrok Lite, so a guide written last year that promises free image or video generation no longer matches what you get. Treat the free plan as a genuinely useful way to try Grok and handle light tasks, not as a permanent substitute for a paid plan if you generate media or run long sessions. ## SuperGrok vs SuperGrok Lite vs SuperGrok Heavy The standalone subscriptions all carry the SuperGrok name, and the differences between them are mostly about usage limits and how many agents you can run, not three completely different products. Here is what each one is for, with prices from [Grok's plans page](https://grok.com/plans) as of July 2026. ### SuperGrok Lite ($10 a Month) The budget tier. It roughly doubles your conversation length over the free plan, adds one AI agent in Expert mode, lets you try image and video creation at lower quality, and lifts your limits at regular speed. It is the right pick if you use Grok often enough to resent the free caps but do not need the heaviest features. ### SuperGrok ($30 a Month) The mainstream plan and the one most paying users want. It adds access to Grok Build, five times longer conversations than the free tier, faster replies, HD video generation, and more file uploads. xAI currently runs a seven-day free trial on it, so you can test it at no cost before the first $30 charge, as long as you cancel in time if it is not for you. ### SuperGrok Heavy ($99 Promo, Then $300 a Month) The power tier, currently discounted to $99 a month for the first three months. It includes everything in SuperGrok plus the highest usage limits, a multi-agent mode that runs up to sixteen agents in parallel on hard problems, dedicated support, and early access to new features. For most people it is overkill. It earns its price only for a narrow group running heavy reasoning or research workloads daily. Some paying users have reported that usage limits and media quotas tightened through mid-2026, with video and voice throttled faster than before. These are user reports rather than published figures, and limits move often, so check the current allowances on your plan before assuming a tier will cover your volume. ## X Premium vs Standalone SuperGrok: Which One Gives You Grok? This is where most people overpay, so it is worth slowing down. There are two separate ways to buy Grok as an individual, and they are easy to confuse because both involve Elon Musk's companies and both put "Grok" on the label. The first is **standalone SuperGrok**, bought at grok.com. You are paying for Grok itself, and nothing else. The second is **X Premium**, the subscription to the X social platform, which bundles some Grok access in alongside ad-free posts, longer videos, and the other X perks. X Premium runs around $8 a month and X Premium+ around $40, though those figures vary by region and have changed over time, so confirm them on X before subscribing. The trap is assuming X Premium+ at roughly $40 gives you the full Grok experience. It does not. The X bundles include Grok with tighter caps and fewer of the standalone features, while a $30 SuperGrok subscription gives you more Grok capability for less money. If you mainly want Grok, the standalone SuperGrok line is almost always the better buy. If you live on X anyway and just want Grok as a bonus, the bundle can make sense. xAI has also at times discounted SuperGrok for existing X subscribers, so if you already pay for X, check whether a current bundle discount changes the math.
You wantBuyRoughly per month
Just to try GrokFree tier at grok.com$0
More Grok, lowest paid priceSuperGrok Lite$10
Grok as your main AI toolSuperGrok$30
X perks plus some GrokX Premium / Premium+$8 / $40
Maximum Grok powerSuperGrok Heavy$99 (promo) / $300
## Grok API Pricing for Developers If you are building software on Grok rather than chatting with it, you do not buy a subscription at all. You pay per [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), the chunks of text the model reads and writes, billed separately for input and output. A SuperGrok plan does not include API usage, and API usage does not need a SuperGrok plan. Here are the current rates from [xAI's pricing docs](https://docs.x.ai/developers/pricing) as of July 2026.
ModelInput (per 1M tokens)Output (per 1M tokens)Context
grok-4.6$2.00$6.00500K
grok-4.5$2.00$6.00500K
grok-4.3$1.25$2.501M
grok-4.20-0309-reasoning$1.25$2.501M
grok-4.20-0309-non-reasoning$1.25$2.501M
grok-4.20-multi-agent-0309$1.25$2.501M
grok-build-0.1$1.00$2.00256K
The new flagship is [grok-4.6](https://geotoolbox.ai/blog/grok-4-6), released August 12, 2026 at the same $2 input and $6 output per million tokens and 500K-token context window as grok-4.5 before it. Watch the context tier, because the headline price is only half the story: xAI's docs list grok-4.6 at $2 input / $6 output for prompts **under 200K tokens**, and **$4 / $12 above that**, with cached input at $0.50 (rising to $1 above 200K). Since the model's selling point is a 500K window, any prompt that actually uses it bills at the higher rate. That is pricier than the outgoing grok-4.3 ($1.25/$2.50, also doubling above 200K), but still well under rivals like Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). grok-4.3 remains the value option, and at $1.25 input and $2.50 output it undercuts the comparable frontier models from OpenAI and Anthropic, whose top tiers run several times higher. That price-per-intelligence is xAI's main pitch to developers right now. A few things to read into that table. The official docs list a tighter set of models than the third-party pricing aggregators do. Older entries like grok-4, grok-4-fast, and grok-3 still show up on other sites at rates like $3 and $15 per million, but they are not in xAI's current model list, so treat them as legacy unless you confirm them directly. There is no grok-5 in the API yet, which lines up with the model still being [in training rather than released](https://geotoolbox.ai/blog/grok-5). Beyond the per-token rate, server-side tools are billed on top: web and X search run $5 per thousand calls, code execution $5, file attachments $10, and collections search $2.50, per xAI's pricing docs. A quirk worth knowing: on the consumer apps you often cannot tell which Grok model answered a given question, because tiers route between models automatically and the interface does not show it. The API is the opposite. It uses explicit model IDs, most of them dated, so if you need to know exactly what you are running, the API is the only place you get a firm answer. ## How to Use Grok (and Where to Access It) You can reach Grok in four places, and which one you pick changes what you pay and what you get. 1. **grok.com (web).** The simplest path. Go to grok.com, sign in with a Google account or email, and start chatting on the free tier. No X account is required, which surprises people who assume Grok is locked to the social platform. 2. **The X app.** Grok is built into X for logged-in users, and your access there depends on whether you have X Premium. This is the bundled path described above. 3. **The standalone Grok app.** xAI publishes a dedicated Grok app for iOS and Android, which uses the same account and subscription as grok.com. 4. **The xAI API.** For developers. You create a key in the xAI developer console and call the models directly, billed per token as covered above. For most people, the answer is grok.com on the web or the standalone app, on the free tier to start. Upgrade to SuperGrok only once you have hit the free limits enough to know you need more. One caveat: availability is not universal. Grok and X Premium are not sold in every country, and prices differ sharply by region, which is why some users abroad cannot subscribe at the US price or at all. If a plan will not let you check out, regional restrictions are the usual reason. ## Is Grok Worth It? Grok vs ChatGPT, Claude, and Gemini on Price On the consumer side, it depends on the tier. SuperGrok Lite at $10 undercuts every $20 rival, but the mainstream SuperGrok tier at $30 sits above ChatGPT Plus, [Claude Pro](https://claude.com/pricing), and [Google's Gemini](https://gemini.google/subscriptions/) AI Pro, which all cluster around $20. So at the tier most people actually buy, Grok is not the cheapest, and reviewers who do not care about real-time X data often struggle to justify the extra $10 over a $20 rival. On the API, the story flips. grok-4.3 at $1.25 input and $2.50 output is one of the cheapest frontier models available, well below the top OpenAI and Anthropic tiers.
ProviderMain paid planFlagship API (input / output per 1M)
xAI GrokSuperGrok $30grok-4.6, $2.00 / $6.00 (grok-4.3, $1.25 / $2.50)
OpenAI ChatGPTChatGPT Plus about $20GPT-5.6 Sol, about $5.00 / $30.00
Anthropic ClaudeClaude Pro about $20Claude Opus 5, about $5.00 / $25.00
Google GeminiGoogle AI Pro $19.99Gemini 3.1 Pro, about $2.00 / $12.00
So whether Grok is worth it depends on which buyer you are. As a $30 consumer chatbot, it is a harder sell than the $20 alternatives unless you specifically want its real-time access to X and its looser content style. As an API for developers, grok-4.3 is genuinely cheap for a frontier model, which is the strongest case for paying xAI anything. And if you live on X already, the bundle math changes again. For a feature-by-feature read rather than just price, see our comparisons of [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt) and [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude), and our standalone breakdown of [Gemini's pricing](https://geotoolbox.ai/blog/gemini-pricing). ## Canceling, Refunds, and Free Trials Cancellation trips people up because it depends on where you paid. If you subscribed on grok.com, you cancel in your account settings on the web. If you subscribed through the Apple App Store or Google Play, you cancel in that store's subscription list, not on grok.com, because the store is what charges you. The quickest way to find out which one applies is to check your card statement or the store you used to sign up. Two practical warnings. First, subscriptions are generally non-refundable once a billing period starts, so canceling stops the next charge rather than refunding the current one. If you want to avoid a charge, cancel before the period renews, not after. Second, users have reported the web cancellation button failing to open when an ad-blocker or browser extension interferes, so if it does not open, try an incognito window. And if you start the SuperGrok seven-day free trial, it converts to the full $30 a month automatically, so set a reminder before day seven if you only wanted to test it. ## Which Grok Plan Should You Actually Pay For? Most people overbuy. Here is how to decide, from the bottom up. **Stay free** if you ask Grok a handful of questions a day and do not generate much media. The free tier covers it, and you will know you have outgrown it the day you keep hitting the limit mid-task. **Pay $10 for SuperGrok Lite** if you use Grok regularly but want the lowest paid price and do not need the heaviest features or high-resolution media. **Pay $30 for SuperGrok** if Grok is one of your main AI tools. This is the right answer for most paying users, and the seven-day trial lets you confirm it costs you nothing to find out. **Pay for SuperGrok Heavy** only if you run heavy reasoning or research daily and will use the multi-agent mode. The $99 promo makes it easier to try, but at the full $300 it is a professional tool, not an everyday one. **Use the API instead of a plan** if you are building software. A subscription buys one person a seat in the app; the API meters whatever your code does, and grok-4.3 is cheap enough that for programmatic work it is usually the right model. ## Why Grok's Price Matters for Your Brand One point here outlasts any specific price. Whichever tier people pay for, Grok is not just a chatbot they visit. It answers questions by pulling in real-time information from X and the web, which means it is already describing your category and naming options to people who ask it what to buy or who to hire. So the more useful question is not which plan you should buy. It is whether Grok mentions your brand at all when someone asks it for a recommendation, and what it says when it does. That visibility does not come on a pricing tier, and it is the gap we help businesses close. The prices above will move again, because xAI has repriced Grok more than once this year, so confirm the current figure on Grok's plans page before you pay. And once you know what Grok costs, the next question is what it tells people about you. Our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) maps where Grok and the other AI engines cite sources your brand is missing from, so you can see which conversations to get into. You can learn the broader method in our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). ## Frequently Asked Questions ### How much does Grok cost per month? Grok is free to start. Standalone paid plans are SuperGrok Lite at $10, SuperGrok at $30, and SuperGrok Heavy at $300, currently discounted to $99 for the first three months, as of July 2026. Team plans are $30 per seat, and the developer API is billed separately per token. ### Is the Grok free tier enough? For light use, yes. The free tier gives you Grok on the web and in the X app with real-time answers from the web and X. Usage is capped, with users reporting roughly ten prompts in a rolling two-hour window, so you will hit the wall quickly on long sessions. Image and video creation through Grok Imagine have moved to the paid tiers. ### How much does Grok 4.3 cost on the API? grok-4.3 costs $1.25 per million input tokens and $2.50 per million output tokens, with a one-million-token context window, per xAI's developer docs. That makes it one of the cheaper frontier models, well below the top OpenAI and Anthropic tiers. ### Is Grok cheaper than ChatGPT? It depends on the tier. SuperGrok Lite at $10 undercuts ChatGPT Plus, but the mainstream SuperGrok tier at $30 costs more than Plus at about $20. On the API it is clearer: grok-4.3 at $1.25 and $2.50 per million tokens is cheaper than [GPT-5.6 Sol](https://developers.openai.com/api/docs/pricing), which runs about $5 and $30. ### Do I need an X account to use Grok? No. You can use Grok at grok.com by signing in with a Google account or email, no X account required. An X Premium subscription is a separate path that bundles some Grok access with the social platform. ### How do I cancel a Grok subscription, and can I get a refund? Cancel from wherever you subscribed: grok.com settings for web sign-ups, or the Apple or Google subscription list if you paid through an app store. Subscriptions are generally non-refundable, so canceling stops the next charge rather than refunding the current period. ## Sources - Grok plans and pricing - Grok (xAI) - `grok.com/plans` - xAI API pricing - xAI Docs - `docs.x.ai/developers/pricing` - New Grok Imagine limits spark user response - Calcalist - `calcalistech.com/ctechnews/article/rkbynj99bx` - Claude plans and pricing - Anthropic - `claude.com/pricing` - Google AI Pro and Ultra subscriptions - Gemini - `gemini.google/subscriptions` - OpenAI API pricing (GPT-5.6 Sol) - OpenAI - `developers.openai.com/api/docs/pricing` --- ## Best Content Optimization Tools (2026): Rank and Get Cited > The best content optimization tools in 2026, tested: what each does, real pricing, who should skip it, and where they fall short on getting cited by AI. - Canonical: https://geotoolbox.ai/blog/best-content-optimization-tools - Published: 2026-06-26 · Updated: 2026-08-22 The best content optimization tools all do the same core thing: they score your draft against the pages already ranking and hand you a target to write toward. The hard part is not finding one. It is knowing what that score is actually worth, which tool fits your budget and workflow, and why a 95 in Surfer can still leave you invisible in the AI answers your buyers now read.
![A 94/100 content-optimizer score next to an AI answer that cites other sources instead.](/blog/best-content-optimization-tools/content-score-is-not-a-citation.png)
A top optimizer score is real work — but it is not the same as earning the AI citation.
This is a tested list of the content optimization tools worth paying for in 2026: what each is best at, the real price, who should skip it, and how each one handles AI citation. > **Disclosure:** GEO Toolbox, the site you are reading this on, makes the Content Analyzer recommended for the citability job later in this guide. The optimizer reviews below are of other companies' products, several linked through affiliate programs as disclosed above. ## What Content Optimization Tools Actually Do A content optimization tool takes your draft and a target keyword, pulls the pages already ranking for that keyword, and scores how well your text matches them. The output is a number, usually 0 to 100, plus a checklist: add these terms, hit this word count, use this many headings. Under the hood they run natural language processing (NLP) over the top results and extract the terms those pages share. Surfer and Frase do this live inside a Google Docs or WordPress editor, so the score moves as you type. That real-time loop is the whole appeal. You stop guessing whether you covered the topic and get a target to write toward. Most tools bundle three jobs that are worth separating: - **Term coverage and scoring**, the core feature, comparing your draft to ranked competitors. - **Briefs and research**, pulling the questions, headings, and entities to include before you write. This is where [content chunking](https://geotoolbox.ai/glossary/content-chunking) and entity coverage get decided. - **AI writing**, generating the draft itself. In practice, the scoring is the part that works, the briefs save real time, and the built-in AI writer is the part most experienced writers turn off. This is the optimization side of the job. The writing-craft side, how to structure sentences and frame facts so a page earns a citation, is a separate skill we cover in [AI content optimization](https://geotoolbox.ai/blog/ai-content-optimization). This guide is about the tools. ## A High Score Is Not a Ranking, and Definitely Not an AI Citation Here is what the roundups leave out. A content score measures how closely your page resembles the pages already ranking. It does not measure whether yours will rank, and the gap between those two things is wide. Surfer's own [study of over a million SERP entries](https://surferseo.com/blog/surfer-content-score-study/) put the correlation between its content score and ranking position at 0.28. That is a weak positive relationship, and Surfer publishes it honestly. An independent [Ahrefs study](https://ahrefs.com/blog/seo-content-score-study/) of several optimization tools found some scores, Frase's among them, correlating closer to 0.1. Different studies, different samples, same lesson: the score clears the first gate, topical relevance, and then real ranking factors take over. Search Engine Land framed it exactly that way in [content scoring works, but only for the first gate](https://searchengineland.com/content-scoring-tools-work-but-only-for-the-first-gate-in-googles-pipeline-469871). Chasing the number past that gate backfires. Push a draft to 100 and you start adding terms the reader does not need and headings that exist only to satisfy the tool. Most practitioners settle around 70 rather than 100, because the last 30 points are usually where natural writing goes to die. In our experience, the pages that get cited are the ones that stopped optimizing at "covered" and spent the rest of the effort on something a tool cannot score. Which is the second blind spot, and it survived the AI gold rush. Most of these tools bolted on an AI-visibility tracker in 2025 and 2026, so they can now tell you whether your brand gets mentioned in AI answers. Useful, but separate. The content score you actually optimize toward is still built from the top-ranking Google pages. A 90-plus score tells you nothing about whether ChatGPT, Perplexity, or Google's AI Overviews will quote the page in front of you, because the signals that earn an AI citation, clear direct answers, original data, and the [trust signals that make content citable](https://geotoolbox.ai/blog/eeat-ai-search), are not what the score measures. ## How We Judged These Tools
![Five numbered criteria cards describing how the content optimization tools were judged.](/blog/best-content-optimization-tools/how-we-judged-tools.png)
The five criteria behind every review below — vendor-checked in June 2026.
Five criteria, in plain terms. What each tool is genuinely best at. The real entry price, checked on the vendor's own pricing page in June 2026 rather than a stale blog figure. Who should not buy it, because every tool is the wrong choice for someone. Whether it does anything for AI citations. And whether the score it hands you is worth chasing. Prices below are monthly in USD where the vendor publishes USD. Several of these tools restructured pricing in late 2025, so confirm the current number on the pricing page before you buy. ## Content Optimization Tools at a Glance The fourth column is where 2026 gets interesting. Most of these tools added some form of AI tracking in the last year, so the old line "they ignore AI" is dead. What separates them now is whether that tracking is built in or a pricey add-on, and whether it actually grades your page or just watches the leaderboard. We pick the list apart below, then deal with the gap that is left. ## Surfer SEO Best for teams and agencies optimizing a lot of content fast. [Surfer SEO](/go/surfer?ref=best-content-optimization-tools-surfer) is the tool most people picture when they hear "content score." Its Content Editor pulls live SERP data, sets a word count and term list, and updates the score as you write, inside the browser or Google Docs. The learning curve is shallow, the interface is clean, and for production teams that value is real. Pricing starts with a Discovery tier around $49 a month billed yearly, with the more capable Standard plan near $99 and Scale and Enterprise above it. Surfer restructured its plans in late 2025, so confirm the current tier names and figures on its pricing page before you commit. The catch is the add-ons. The headline price covers the editor, but the AI writer and the separate AI tracker are metered or gated to higher tiers, and a few AI-written articles can quietly cost more than the subscription itself. Surfer's own blog admits the generated drafts can [feel like AI](https://surferseo.com/blog/ai-generated-content/), which is a refreshingly honest thing for a vendor to say and a good reason to use the editor and skip the writer. **Who should skip it.** Solo bloggers writing a handful of posts a month will not use enough of Surfer to justify the price. A budget tool like NeuronWriter covers the same core scoring job for a fraction of the cost. One thing to know before committing: cancel, and Surfer deletes your saved editors, audits, and tracked sites at the end of the billing period. Surfer shipped an AI Tracker in 2025 that monitors your brand's mentions across ChatGPT, Perplexity, and Google's AI Overviews, sold as a paid add-on in the higher tiers. It watches the leaderboard rather than grading a page. It can tell you that you are not showing up; it will not grade the page in front of you for why, and the content score you optimize toward is still pure Google SERP. ## Frase Best for writers who want a brief and a draft in one place without paying enterprise prices. [Frase](/go/frase?ref=best-content-optimization-tools-frase) starts from the SERP, builds an outline and a question list in minutes, and folds optimization scoring into the same editor. For a freelancer or small team, that research-to-draft loop is the selling point, and it is genuinely fast. Pricing is the friendliest of the serious tools. The Starter plan is $49 a month and the Professional plan $129, with a 7-day trial and a discount on annual billing, per the [Frase pricing page](https://www.frase.io/pricing). That puts a capable optimizer within reach of a solo budget. The real weakness is the score. The independent Ahrefs study mentioned earlier put Frase's content score among the weakest correlators with actual rankings, and reviewers consistently note its scoring is less predictive than Surfer's. Treat Frase's number as a coverage check, not a ranking forecast. Its AI writer, like the others, produces drafts that need real editing before they publish. **Who should skip it.** If your team lives by the optimization score and wants the tightest correlation with rankings, Frase will frustrate you. It wins on workflow and price, not on scoring precision. For AI citation, Frase has gone furthest of the optimizers. Its 2026 relaunch built AI search tracking into every plan, covering ChatGPT, Claude, Gemini, Perplexity, and AI Overviews, and added a separate GEO score that grades pages and lists fixes, measured against what AI actually cites rather than the top Google results. If you want a content editor that takes AI citation seriously, Frase is the one to beat. The catch is that all of it lives inside a paid Frase plan and its editor workflow. ## Scalenut Best for teams that want planning, writing, and optimization in one fast workflow. [Scalenut](/go/scalenut?ref=best-content-optimization-tools-scalenut) covers the whole content lifecycle, from keyword clusters to a drafted, optimized article, and its Cruise Mode can take a keyword to a full draft in a few minutes. If speed and a single interface matter more than best-in-class depth, it fits. Pricing is mid-budget. The Starter plan is $59 a month, Plus is $89, and Professional is $199, with a 7-day trial, per the [Scalenut pricing page](https://www.scalenut.com/pricing). Promotions run often, but the list prices are what you should plan around. The weakness is depth and output quality. Scalenut's SERP and NLP analysis is solid but lags Surfer and Clearscope for teams that want the most precise scoring, and its AI drafts trend generic and repetitive, which means more editing than the "minutes to a draft" pitch implies. Some users also report the app feeling slow under load. **Who should skip it.** If you care most about the quality of the generated draft, no all-in-one tool will satisfy you, and Scalenut is no exception. Use it for the speed of the workflow, not for hands-off writing. Scalenut added AI-visibility tracking in 2025, monitoring your mentions across ChatGPT, Perplexity, Gemini, and AI Overviews, plus some page-level AI recommendations. It is real tracking, built into the plans rather than sold as a separate add-on. What it does not give you is a clean letter grade for a specific live page's citability. ## NeuronWriter Best for budget-conscious bloggers and SEOs who want real semantic optimization without a big-team price tag. [NeuronWriter](/go/neuronwriter?ref=best-content-optimization-tools-neuronwriter) does the core job, SERP analysis, term suggestions, internal-link recommendations, and a content score, for the lowest entry price on this list. That price is the headline. The Bronze plan is about $23 a month, or $19 on annual billing, with Silver and Gold tiers stepping up from there. For a writer publishing a few posts a month, NeuronWriter optimizes each article for roughly a dollar, which is why it is the value pick here. The trade-off is polish. The interface is more utilitarian than Surfer's or Clearscope's, support is lighter, and the built-in AI writer produces the weakest drafts of the group. You are buying the scoring and the brief, not the writing. **Who should skip it.** Teams that need a refined interface, onboarding, and a dedicated account manager will find NeuronWriter bare. It is a solo and small-shop tool, and it is honest about being one. NeuronWriter added an AI visibility module in 2026 that tracks mentions in ChatGPT, Perplexity, and AI Overviews, though it can require a separate add-on. Like the others at this price, it monitors rather than grades. As a low-cost way to make sure a draft covers its topic, it is still hard to beat. ## Clearscope Best for in-house editorial teams and premium publishers that put content quality first. Clearscope built its reputation on a clean grading algorithm and tight Google Docs integration, and its term recommendations are among the most trusted in the category. If your job is to make good writers slightly better and hold a house standard, it is excellent. A note on honesty: Clearscope has no affiliate program, so the link here earns us nothing. We include it because leaving out the best premium grader would make this list less useful, which is the whole point. Price is the barrier. Clearscope starts at $129 a month for Essentials and jumps to $399 for Business, with no free trial, per its [pricing page](https://www.clearscope.io/pricing). That is built for funded teams, not solo writers, and it is the most common complaint in reviews. The tool is also narrow by design. It grades and recommends; it does not plan clusters, audit a whole site, or monitor content decay the way MarketMuse does. You are paying a premium for one job done very well. **Who should skip it.** Solos and small teams. The per-report cost is hard to justify below a steady publishing volume, and Frase or NeuronWriter cover the core scoring for far less. Clearscope added prompt tracking that watches whether ChatGPT and Gemini mention your brand. Worth knowing: its well-known A-to-F grade is an SEO content grade, not an AI-citation grade. The tracking tells you where you stand; the grade still measures resemblance to top Google results. ## MarketMuse Best for teams building topical authority across a whole site rather than optimizing one page at a time. MarketMuse models a topic in depth, finds content gaps, plans clusters, and is one of the few tools with native content-decay monitoring, flagging pages that are slipping before you notice the traffic drop. It runs a free plan plus paid tiers above it. MarketMuse no longer publishes its prices and routes you to a demo, so confirm the figure for your team size on the [MarketMuse pricing page](https://www.marketmuse.com/pricing). It has no standard affiliate program either, so this is a straight recommendation, not a paid one. The weakness is the learning curve. MarketMuse is a strategy platform, and for a small site without a cluster plan it is overkill, with a feature set you will pay for and barely use. Its AI first drafts, like most on this list, need heavy editing. **Who should skip it.** Single-page optimizers and small blogs. If you are not running a site-wide content strategy, the depth is wasted and a simpler editor will serve you better and cheaper. On AI citation, MarketMuse is the holdout. Since the Siteimprove acquisition it has added no AI-visibility tracking, so it is the one tool here that still cannot tell you whether AI mentions you. It stays a planning brain: thorough topical coverage helps AI pick sources, but MarketMuse measures none of it. ## SE Ranking Best for people who want content optimization without buying a separate tool for it. [SE Ranking](/go/seranking?ref=best-content-optimization-tools-seranking) is a full SEO platform, rank tracking, audits, and keyword research, and it folds a content editor and optimization scoring into the same subscription. If you are already shopping for an all-in-one and want the writing-time optimizer included, it consolidates the stack. Pricing reflects the broader scope. Plans start around $90 a month on annual billing and more month to month, after a late-2025 restructure into Core, Growth, and Enterprise tiers, with a 14-day trial, per the [SE Ranking pricing page](https://seranking.com/subscription.html). You are paying for the whole platform, not just the optimizer. **Who should skip it.** If all you want is a content editor, a dedicated tool like Surfer or NeuronWriter does that one job better, and you will not touch most of SE Ranking's other modules. SE Ranking also runs an AI Visibility Tracker, monitoring AI Overviews, AI Mode, ChatGPT, Gemini, and Perplexity, sold as an add-on or standalone on top of the base plan. Like most here it tracks mentions rather than grading a page, but it is one of the more complete monitors if you want AI visibility inside a single SEO platform. ## A Few Others Worth Naming Four tools did not earn a full entry but belong on your radar. Semrush bundles a content optimizer, the SEO Writing Assistant, inside its larger suite; it is convenient if you already pay for Semrush, but weaker than a dedicated scorer. Dashword is a lighter, cheaper brief-and-score tool that small teams like. PageOptimizer Pro is the technical, manual option favored by SEOs who want to control the inputs. GrowthBar folds optimization into a blogging workflow. None changes the picture above, and the same caveat applies to all of them: the score is a Google-resemblance metric, not a citability grade. ## What the AI Trackers Still Don't Do: Grade Your Live Page The AI gold rush reached the optimizers, and most of them now track whether AI engines mention your brand. That closes one gap and leaves another wide open. Tracking tells you where you stand. It does not tell you why a specific page is or is not getting cited, or what to change. And almost every AI feature above lives inside a paid writing editor and applies to content you draft there, not to a live URL you already published. Getting your existing pages quoted is its own discipline, [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), and it is real enough that researchers now build [formal evaluations for it](https://arxiv.org/abs/2602.12187). Frase goes furthest among the optimizers, with a GEO score that grades pages and lists fixes, all inside its paid platform. So the real split is not two buckets, it is three jobs.
JobWhat it answersTools that do it
Optimize for GoogleDoes my draft resemble the ranking pages?All seven optimizers
Track AI mentionsAm I showing up in AI answers, and where?Surfer, Frase, Clearscope, SE Ranking, Scalenut, NeuronWriter
Grade a live page for citabilityWhy is this page not cited, and what do I fix?Frase (in its editor), GEO Toolbox (any URL)
That third job is the one the category still under-serves, and the dedicated [AI-visibility platforms](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools) handle the tracking without the page-level grade either. [Our Content Analyzer](https://geotoolbox.ai/features/content-analyzer) is built for it: paste any live URL and it returns an A-to-F citability grade, a check of which AI engines cite the page, and the specific fixes, scored against the signals shared by the pages AI actually cites, the direct answers, the structured facts, the [entity coverage](https://geotoolbox.ai/blog/entity-seo) and trust markers, not resemblance to the top ten Google results. It works on any URL you paste, not just on content you draft inside an editor, and it is included on every GEO Toolbox plan, [from $99 a month](https://geotoolbox.ai/pricing) ($79 billed annually) with a 7-day free trial — in the same price band as the optimizers above. Pair an optimizer with a real citability grade and you cover all three jobs: rank with Surfer or Frase, and find out whether your best-optimized page is actually citable or just well-scored. ## Which Content Optimization Tool Should You Actually Buy? The right answer depends on the job, not the leaderboard. Mapped to the most common situations: - **Solo blogger or freelancer on a budget:** [NeuronWriter](/go/neuronwriter?ref=best-content-optimization-tools-neuronwriter) for the cheapest real optimization, or [Frase](/go/frase?ref=best-content-optimization-tools-frase) if you also want briefs and a draft in one place. - **Content team optimizing at volume:** [Surfer SEO](/go/surfer?ref=best-content-optimization-tools-surfer). The editor is fast, the workflow scales, and the score is the most ranking-correlated of the group. - **Editorial team that lives on quality:** Clearscope, if the budget is there and grading is the priority. - **Site-wide content strategy:** MarketMuse for clusters, gaps, and decay monitoring. - **One tool for everything:** [SE Ranking](/go/seranking?ref=best-content-optimization-tools-seranking) or [Scalenut](/go/scalenut?ref=best-content-optimization-tools-scalenut), depending on whether you want a full SEO suite or a content-first lifecycle. - **Brand that needs to show up in AI answers:** any optimizer above for Google, paired with [a tool that measures AI citation](https://geotoolbox.ai/blog/how-to-track-ai-visibility). Whichever you pick, set the score target at "covered" rather than 100, and spend the time you save on original data and clear answers. That is the part that ranks past the first gate, and the part an AI engine can actually quote. ## Frequently Asked Questions ### What are content optimization tools? Content optimization tools score a draft against the pages already ranking for your target keyword and tell you which terms, headings, and length to match. Most run inside a Google Docs or WordPress editor and update the score in real time as you write. The best known are Surfer, Frase, and Clearscope. ### Do content optimization tools actually improve rankings? They help with one part of ranking, topical relevance, and not much beyond it. Surfer's own study found its content score correlates with ranking position at about 0.28, a weak positive relationship, and independent tests put some tools lower. The score clears Google's first gate; links, authority, and intent decide the rest. ### What is the best free content optimization tool? There is no strong fully free option, but the cheapest paid tools cover the core job. NeuronWriter starts around $23 a month ($19 on annual billing), and MarketMuse has a limited free plan. Frase offers a 7-day trial rather than a free tier. For a genuinely free check, Mangools runs a single-page content optimizer with no account needed. ### Surfer vs Frase: which should I pick? Pick Surfer if you optimize at volume and want the most ranking-correlated score in a fast editor. Pick Frase if you want SERP research, a brief, and a draft in one cheaper interface. Surfer wins on scoring precision; Frase wins on workflow and price. ### Do content optimization tools help you get cited by ChatGPT? Several now track it. Surfer, Frase, Clearscope, SE Ranking, and Scalenut all added AI-visibility tracking in 2025 and 2026, so they can show whether ChatGPT, Perplexity, or AI Overviews mention your brand. Tracking is not the same as fixing, though. Frase comes closest, with GEO scoring inside its editor. What stays uncommon is a tool built to grade any URL you paste for citability and hand you the fixes, separate from a Google content score with tracking bolted on. ### Can you over-optimize content? Yes. Pushing a draft to a 100 score usually means adding terms and headings the reader does not need, which reads worse and can do more harm than good. Most experienced writers stop around 70 and spend the rest of the effort on originality and clarity. ## The Score Is a Means, Not the Goal A content optimization tool is worth paying for. It turns "did I cover the topic" from a guess into a target, and on a team it keeps quality consistent. Just hold the number in perspective: it confirms you cleared topical relevance, not that you will rank, and not that any AI engine will quote you. Pick the tool that fits your budget and workflow from the list above, optimize to "covered," and put the saved hours into original data and clear answers. Then check the part the optimizers track but do not grade: [grade any live URL for AI-citability](https://geotoolbox.ai/features/content-analyzer) with our Content Analyzer, and find out whether your best-optimized page is citable or just well-scored. ## Sources - Surfer SEO content score study, content score to ranking-position correlation (0.28) - `surferseo.com/blog/surfer-content-score-study` - Ahrefs SEO content score study, independent multi-tool score-versus-ranking analysis - `ahrefs.com/blog/seo-content-score-study` - Search Engine Land: content scoring works, but only for the first gate - `searchengineland.com/content-scoring-tools-work-but-only-for-the-first-gate-in-googles-pipeline-469871` - Surfer SEO: making AI-generated content not feel like AI - `surferseo.com/blog/ai-generated-content` - SAGEO Arena: evaluating generative engine optimization (arXiv) - `arxiv.org/abs/2602.12187` - Pricing pages, verified June 2026: Frase, Scalenut, NeuronWriter, Clearscope, MarketMuse, SE Ranking - `frase.io/pricing` - `scalenut.com/pricing` - `neuronwriter.com/pricing` - `clearscope.io/pricing` - `marketmuse.com/pricing` - `seranking.com/subscription.html` --- ## Best Link Building Platforms (2026): The 5 Types, by Job > The best link building platforms in 2026 by type: research tools, outreach software, marketplaces, and managed PR. Who to skip, and what earns AI citations. - Canonical: https://geotoolbox.ai/blog/best-link-building-platforms - Published: 2026-06-26 · Updated: 2026-08-22 A link building platform is any software that helps you earn, find, or buy backlinks. That one definition covers at least five very different kinds of product, which is why most "best link building platforms" lists are close to useless. They rank a backlink research tool against a guest-post marketplace against a managed agency, as if you would ever choose between them. You would not. You pick the category that fits the job you are stuck on, then the tool inside it. This guide sorts the platforms by type, names the ones worth paying for in 2026, and answers the question every roundup skips: do the links you build still matter now that your buyers read AI answers instead of scrolling ten blue links? They do. The goal just moved. ## What a Link Building Platform Actually Is The term covers five categories. They solve different parts of the job and are often combined, but the tool that is best at one is rarely the one you want for another. 1. **Backlink research and monitoring tools** show you who links to any site, your own or a competitor's. This is the analysis layer, not acquisition. Ahrefs and Semrush live here. 2. **Outreach and prospecting platforms** run the manual work: finding the right contact and a verified address, then managing email sequences at scale. Hunter and Snov.io handle the contact data; Pitchbox, Respona, and BuzzStream run the campaigns. The two layers are usually stacked. 3. **Link marketplaces** sell placements directly. You browse publishers, filter by metrics, and pay for a post. Collaborator, Adsy, and Getfluence work this way. 4. **Journalist and reactive-PR platforms** connect you to reporters who need a source, so you earn an editorial mention by answering a query. Qwoted, Featured, and the revived HARO sit here. This is the affordable route to the press mentions that matter most for AI visibility. 5. **Managed services and digital PR** do the work for you, from a single placement to a full campaign that earns coverage in real publications. The cost, the control, and the penalty risk change sharply as you move through that list. A research subscription is cheap and risk-free. A marketplace is fast but puts the quality bar entirely on you. The table below shows the split.
Platform type What it does Best for Rough cost Penalty risk
Research and monitoring Audit any site's backlinks; find prospects and broken links Strategy, competitor analysis, tracking ~$100-200/mo None
Outreach automation Run and track email campaigns at scale Teams sending hundreds of pitches ~$50-500/mo None (you still earn the link)
Prospecting and email data Find the right contact and a verified email Solo SEOs and small outreach teams ~$30-100/mo None
Link marketplaces Buy placements from listed publishers Speed and volume, when you can vet quality ~$50-500 per link Medium to high
Journalist / reactive PR Answer reporter queries to earn editorial mentions Earning press mentions on a budget Free to ~$100/mo None
Managed services and digital PR Done-for-you outreach, placements, or press Brands with budget and no time ~$1,500-10,000/mo Low to high (depends on the vendor)
Read that table as a decision tree, not a ranking. The rest of this guide takes each category in turn, but first, the question that decides whether any of this is worth doing in 2026. ## Do Backlinks Still Matter for AI Search? Yes, but the reason changed, and that change should steer which platform you pick.
![Bar chart comparing brand mentions and backlinks correlation with AI Overviews visibility.](/blog/best-link-building-platforms/mentions-vs-backlinks-correlation.png)
In Ahrefs's 75,000-brand study, web mentions correlated with AI Overviews presence roughly three times as strongly as backlink count.
For Google's classic ranking, links are still foundational. A page with relevant, trusted backlinks tends to rank better. For AI search, the link matters less directly. Backlinks help your page get discovered and indexed, which is the precondition for an engine retrieving it at all, but they are not the thing that gets you quoted. That is a step beyond the link itself. The data is worth being precise about. Ahrefs studied 75,000 brands and measured what correlated with showing up in Google's AI Overviews. Web mentions of the brand correlated at 0.664; backlinks at 0.218, per the [Ahrefs brand-mentions study](https://ahrefs.com/blog/ai-overview-brand-correlation/). Both are correlations, not proof of cause, and Ahrefs says as much. But the gap is wide: being mentioned, named, and discussed across the web correlated with AI visibility roughly three times as strongly as the raw backlink count. Taken alone, that means links by themselves barely move AI citations, which is true. What moves them is the work behind a good link: the relevance, the relationships, the reason a real site chose to link to you in the first place. Semrush's [study of the most-cited domains in AI answers](https://www.semrush.com/blog/most-cited-domains-ai/) points the same way, with community and editorial sources like Reddit ranking among the most-cited domains across engines. That is good news for anyone still building links. The work that earns a strong editorial link (a useful resource, original data, a genuine relationship with a publisher) is the same work that earns a [brand mention](https://geotoolbox.ai/glossary/brand-mention) and an AI citation. A guest post on a relevant industry site builds a link and puts your name in front of the model that later answers a buyer's question. A digital PR campaign earns a backlink and a sentence of context that an engine can quote. So the platform you choose should be judged on whether it produces that kind of link, the kind tied to a real publisher and a real audience, or just a URL on a page nobody reads. The cheap end of the marketplace world fails that test. Digital PR passes it easily. Most of this guide is about telling those apart, because the [trust signals that make content citable](https://geotoolbox.ai/blog/eeat-ai-search) and the [entity signals AI engines use to understand your brand](https://geotoolbox.ai/blog/entity-seo) are downstream of the same relevance-first link work, not separate from it. ## Backlink Research and Monitoring Tools Start here even if you never plan to buy a link. Research tools tell you who links to your competitors, which of their pages earn the most links, and what is worth chasing. Every other category works better when you point it at targets these tools surfaced. Ahrefs and Semrush are the two that matter. Ahrefs is the backlink specialist of the pair, with one of the freshest, broadest indexes and the cleanest workflow for finding prospects: pull a competitor's referring domains, filter by traffic and relevance, and export a target list. Its [link building guide](https://ahrefs.com/seo/link-building) is also the reference most practitioners learned from. Semrush bundles backlink data into a wider suite, so if you already pay for it for keywords and rank tracking, its Backlink Gap and Link Building tools cover most of the same ground without a second subscription. For a full head-to-head on their keyword and backlink data using real first-party numbers, see our [Semrush vs Ahrefs comparison](https://geotoolbox.ai/blog/semrush-vs-ahrefs). This category also covers backlink monitoring tools, the narrower job of watching your own profile: which links you gained, which went missing, and whether a toxic batch appeared overnight. Most monitoring lives inside Ahrefs and Semrush, with cheaper standalone options like Linkody and Monitor Backlinks for teams that only need the watch function. Skip a dedicated research tool if you are buying a handful of links a year through a managed service: a $100-plus monthly subscription to audit them is overkill. Ask the vendor for the metrics and spend the money on the placement instead. Research tools earn their keep when you run outreach yourself and need a steady supply of vetted prospects. ## Outreach and Prospecting Platforms This is where most real link building happens: you find sites worth a link, identify the right person, and pitch them something they actually want. Two layers of software support it. The outreach layer manages campaigns. [Pitchbox](https://pitchbox.com/) is the enterprise standard, built for agencies running hundreds of sequences with team reporting, and priced accordingly. Respona pairs outreach with built-in prospecting and is the common pick for content-led teams. BuzzStream is the long-running option for smaller teams that want a CRM for relationships without enterprise pricing. None of these are wrong; they differ mostly on scale and budget. For teams that want outreach folded into a broader sales-style cadence, [Postaga](/go/postaga?ref=best-link-building-platforms-postaga) and [Instantly](/go/instantly?ref=best-link-building-platforms-instantly) handle campaign automation and deliverability at a lower entry price. The prospecting layer finds the contact. [Hunter](/go/hunter?ref=best-link-building-platforms-hunter) is the best-known email finder: drop in a domain and it returns verified addresses with a confidence score. [Snov.io](/go/snov?ref=best-link-building-platforms-snov) does the same job and adds a built-in drip sequencer, so a solo SEO can prospect and send from one tool instead of stitching two together. Cold outreach reply rates are low, and no tool changes that. A well-targeted, genuinely useful pitch to a relevant site might land a handful of links per hundred sends, and a generic one lands close to zero. Software does not fix a weak pitch or an irrelevant target; it just lets you send more of them faster. That is exactly why the research step comes first. **Who should skip the outreach stack.** If you cannot commit time to writing pitches and following up, do not buy outreach software, because it only multiplies effort you are putting in. A marketplace or a managed service will get you links with far less of your own labor. Outreach platforms pay off for teams that treat link building as an ongoing process, not a one-time purchase. ## Link Marketplaces A link marketplace is a catalog of publishers willing to host a paid post. You filter by domain rating, topic, traffic, and price, place an order, and a placement appears. It is the fastest way to get a link and the easiest place to waste money. [Collaborator](/go/collaborator?ref=best-link-building-platforms-collaborator) is one of the larger transparent marketplaces, with metrics shown per publisher and traffic data attached to each listing. [Adsy](/go/adsy?ref=best-link-building-platforms-adsy) and [Getfluence](/go/getfluence?ref=best-link-building-platforms-getfluence) cover overlapping inventory, with Getfluence leaning toward larger, premium media placements. WhitePress rounds out the category. Most marketplaces sell two products: a fresh guest post written around your link, or a niche edit, also called a link insertion, where your link is added to an article already published. Niche edits look more natural but give you less say over the surrounding context. Either way, the platform is just software; the quality of any single link is on you. The numbers explain why vetting is not optional. A study of more than 400,000 link-selling sites by [Link-Finder](https://link-finder.net/link-building-services/) put the global median at $180 per backlink, with a wide spread by market: a French site runs around $87, while a German one is closer to $430. The same study found that 42.2% of the sites selling links had zero organic traffic, and only 6.5% had a domain rating above 60. Most of what a marketplace will sell you sits on a site few real people visit, which does little for ranking and less for AI citation. Quality is only one of two axes, though, and the second one catches people out. Buying a link that passes ranking signal violates [Google's link spam policy](https://developers.google.com/search/docs/essentials/spam-policies) regardless of how good the site is. A paid dofollow link from a high-traffic, perfectly relevant publisher is still a policy violation; the compliant version carries a `rel="sponsored"` or `rel="nofollow"` attribute and is not meant to pass ranking credit. Marketplaces sell dofollow placements anyway, which is the risk you accept when you buy one. A single bought link is already a violation, but what tends to trigger an actual penalty is volume and velocity: bulk dofollow orders month after month, over-optimized anchor text, or links from a private blog network (PBN), the highest-risk source of all. A few placements from genuinely relevant, trafficked sites are a milder risk than that, but they are not a compliant one, and it is worth being honest with yourself about which you are doing. **Who should skip marketplaces.** If you cannot read a backlink profile and judge whether a site's traffic and metrics are real, do not buy from a marketplace yet, because you will overpay for links that hurt you. Learn to vet first (the checklist is further down), or buy through a managed service that vets for you. ## Reactive PR and Journalist-Request Platforms This is the most underused category, and the cheapest way to earn the kind of editorial mention the AI data rewards. Journalist-request platforms send out reporter queries; you answer with a quote, and if the journalist uses it, you earn a mention and often a link in a real publication. No payment to the publisher, no policy risk, and the placement is genuinely editorial. The landscape shifted recently, so the names matter. HARO, the original, was folded into Cision's Connectively and then discontinued in December 2024; a version reopened under the HARO name in 2025, though many users report it is now flooded with AI-generated answers and thin on quality control. The platforms worth your time now are [Qwoted](https://www.qwoted.com/), which verifies both journalists and sources, and [Featured](https://featured.com/) (formerly Terkel), which runs structured expert questions. Source of Sources is a free option that filled part of the HARO gap. The catch is effort, not money. Winning a placement means writing a genuinely useful, specific answer fast, and most of your responses will go unused. But a single mention in a publication a reporter trusts does more for both your authority and your AI visibility than a stack of bought links, which is why this category punches well above its price. This category is wrong for you if you cannot turn around a sharp, quotable answer within a few hours of a query landing, because the response rate will frustrate you. It rewards subject-matter experts who can write, not teams that want to outsource and forget. ## Managed Services and Digital PR When you have budget but no time, you outsource the whole process. This category runs from white-label link building services that resell placements to full digital PR agencies that earn coverage in real publications. The price range is wide, and so is the quality, so the same vetting rules apply, just aimed at the vendor instead of the link. Two sub-types are worth separating. Blogger outreach and guest posting services handle volume: they pitch and place posts on relevant blogs, often reselling marketplace inventory with a markup. Digital PR is the higher end, and it is the category that matters most for AI visibility. A PR campaign earns editorial mentions in the press, the kind of trusted, talked-about coverage that correlated most strongly with showing up in AI answers in the Ahrefs data above. Tools like [Brand24](/go/brand24?ref=best-link-building-platforms-brand24) and [Prowly](/go/prowly?ref=best-link-building-platforms-prowly) support that work by tracking mentions and managing media relationships, so you can see which placements actually generated coverage and conversation. The trade-off between doing this yourself and hiring out is the same one we cover in [hire an agency or use a tool](https://geotoolbox.ai/blog/geo-services-vs-software). A managed service buys back your time and, with a good vendor, applies vetting you might not have. A bad one resells the same zero-traffic marketplace links at triple the price and reports on domain rating alone. **Who should skip managed services.** If your budget is under a few hundred dollars a month, you cannot buy quality digital PR, and the cheap "managed" packages at that price are usually bulk marketplace links with a dashboard. Either build links yourself with the outreach stack above, or save until you can afford a vendor whose placements you would be proud to show a client. ## How to Vet a Link Building Platform Before You Pay Whatever category you choose, the same checks separate a link worth having from one that wastes money or earns a penalty. Run a site or a vendor through these before you spend. 1. **Check organic traffic, not just domain rating.** Domain rating is easy to inflate. Pull the site in a research tool and confirm it gets steady organic search traffic to relevant pages. The Link-Finder data found 42% of link-selling sites have none, and a site no one visits is rarely worth a link whatever its rating says. The exception is a genuinely authoritative low-traffic site, a trade body or a university page, where relevance carries the value. 2. **Confirm topical and audience relevance.** A link from a site your buyers might actually read carries more weight, for both Google and AI engines, than a high-rating link from an unrelated niche. Relevance is the signal that builds [topical authority](https://geotoolbox.ai/glossary/topical-authority). 3. **Cross-check the metrics for manipulation.** Compare the domain rating against an independent measure like Majestic's Trust Flow versus Citation Flow, and look at the ratio of referring domains to total backlinks. A site with a strong score but a thin, spiky, or redirect-fed link profile has been gamed. Genuine authority looks gradual and broad. 4. **Demand placement proof and reporting.** A credible vendor shows you the live URLs. If a service guarantees rankings, will not name the sites, or reports only "DR 50+ links delivered," walk away. 5. **Watch the link attributes and velocity.** Know whether links are dofollow or carry `rel="sponsored"`, and avoid any pattern that adds dozens of dofollow links a week. Sudden velocity from low-quality sources is the classic penalty trigger. In our experience, the platforms that survive this checklist are the ones backed by a trafficked site and a genuine readership. The ones that fail it are usually the cheapest, the fastest, and the most aggressively marketed, which is exactly why the checklist exists. ## Which Platform Should You Actually Use? Pick by the bottleneck you are stuck at. Which tool tops a generic list is beside the point. **You need to find prospects and contacts.** Start with a research tool to build target lists, then a prospecting tool like [Hunter](/go/hunter?ref=best-link-building-platforms-hunter) or [Snov.io](/go/snov?ref=best-link-building-platforms-snov) for verified emails. This is the cheapest stack and the right starting point for most in-house SEOs. **You are sending pitches and losing track.** Add outreach software. Solo or small team, [BuzzStream](https://buzzstream.com/) or [Postaga](/go/postaga?ref=best-link-building-platforms-postaga); agency scale, Pitchbox or Respona. The tool is for managing volume, so only buy it once you have the volume to manage. **You want links without the labor.** Go to a marketplace like [Collaborator](/go/collaborator?ref=best-link-building-platforms-collaborator) for listings you can filter and vet yourself, or [Adsy](/go/adsy?ref=best-link-building-platforms-adsy) for broader inventory, and run every one through the checklist above. Speed in exchange for doing your own quality control. **You want authority and AI citations.** Start with journalist platforms like Qwoted or Featured for editorial mentions on a budget, then move up to digital PR when you can fund it. These are the slowest routes and the ones most likely to earn the brand mentions that get you quoted in AI answers. They are also, telling enough, the categories with the least software to sell you: the highest-value link work is mostly skill and relationships, not a subscription, which is why the affiliate-heavy roundups tend to skip past it. Most teams end up with a small stack across two categories: a research tool plus one acquisition method that fits their time and budget. Resist buying one of everything. The same discipline applies to the writing side, where a focused set of [content optimization tools](https://geotoolbox.ai/blog/best-content-optimization-tools) beats a drawer full of overlapping subscriptions. ## Measuring Whether Your Links Earn AI Citations Here is the gap most link-building platforms still leave open. They help you build or buy links, and a research tool will tell you the links exist. What they will not show you is whether that work changed what AI engines actually say about your brand. The bigger suites have started adding AI-visibility tracking, but tying a specific link campaign to a movement in AI answers is still mostly guesswork. That is the metric that now matters. If the point of a relevant, trusted link is to get your brand into the answers ChatGPT, Perplexity, and Google's AI Overviews hand your buyers, then the result you track is your [AI citation](https://geotoolbox.ai/glossary/ai-citation) rate, and learning [how to track AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) is its own job, separate from building the links. A campaign that adds twenty links but never moves your presence in AI answers spent its budget on the wrong kind of link. This is the loop geotoolbox is built around. We track where your brand shows up across the major AI engines, so you can watch a digital PR push or a batch of placements against a real change in [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) instead of guessing. It is the same reason a focused set of [GEO tools](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools) belongs next to your link building stack: building the links is half the job, and proving they earned you citations is the half everyone skips. ## The Short Version Stop shopping for "the best link building platform." Decide which job you are stuck on, pick the category that fits it, and judge any tool, marketplace, or agency on whether it produces links a real audience actually sees. That is the kind of link most likely to support a Google ranking and an AI citation alike. Then close the loop. If you are spending on links to win the AI answers your buyers now read, measure that. See where your brand currently shows up across the engines with the free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness), then track whether your link work moves the needle in [geotoolbox's Domain Overview](https://geotoolbox.ai/features/domain-overview). Building links you cannot measure is how budgets disappear; building links you can tie to citations is how you prove the work. ## Frequently Asked Questions ### Are link building services worth it in 2026, or am I just paying for garbage? They are worth it only if you vet what you buy. The Link-Finder study found 42% of link-selling sites have zero organic traffic, so most cheap packages are close to worthless. A handful of relevant, trafficked placements beats a bulk order of fifty every time, and that is the difference between a service worth paying for and one to avoid. ### How much should I pay for a backlink? The global median is around $180 per link, with a wide spread by market and quality. Price should track the site's real organic traffic and topical relevance, not its domain rating, which is easy to inflate. If a link costs $30 and the site has no traffic, you are paying for a number on a dashboard. ### Is buying backlinks against Google's guidelines? Buying links that pass ranking signal violates Google's link spam policy, and paid placements are supposed to carry a `rel="sponsored"` or `rel="nofollow"` attribute. Enforcement is uneven, but the risk is real. Bulk-buying dofollow links at high velocity is the pattern that gets sites penalized, not the occasional vetted placement. ### Do backlinks still matter for AI search like ChatGPT and Perplexity? Yes, but indirectly. Backlinks help your page get discovered and indexed, while brand mentions correlated about three times as strongly as backlinks with showing up in Google's AI Overviews in Ahrefs's 75,000-brand study. The same relevance-first work earns both the link and the citation. ### What is the difference between a link building tool, a marketplace, and a service? A tool helps you find, manage, or analyze links yourself. A marketplace sells you placements you buy directly from listed publishers. A service does the whole job for you, from outreach to placement. Cost and hands-off convenience rise as you move from tool to marketplace to service. ### How can I tell if a backlink is good quality before I buy it? Check that the site gets real organic search traffic to relevant pages, confirm topical relevance to your niche, and cross-check its domain rating against trust and citation flow for signs of manipulation. Then demand a live placement URL, not just a promised metric. No traffic or no transparency means skip it. ## Sources - Ahrefs, brand mentions vs backlinks correlation with AI Overviews (75,000 brands) - `ahrefs.com/blog/ai-overview-brand-correlation` - Ahrefs, Link Building for SEO: The Beginner's Guide - `ahrefs.com/seo/link-building` - Google Search Central, spam policies (link spam and paid links) - `developers.google.com/search/docs/essentials/spam-policies` - Semrush, the most-cited domains in AI answers - `semrush.com/blog/most-cited-domains-ai` - Link-Finder, backlink pricing study (400,000+ sites) - `link-finder.net/link-building-services` --- ## Gemini Pricing in 2026: Plans, API Costs, and Is It Free? > Gemini pricing for August 2026: the free tier, Google AI Plus ($4.99), Pro ($19.99), Ultra ($99.99-$199.99), and the Gemini API token cost by model. - Canonical: https://geotoolbox.ai/blog/gemini-pricing - Published: 2026-06-26 · Updated: 2026-08-18 Google Gemini is free to start, and for a lot of people that is the end of the story. If you want more, the paid consumer plans run $4.99 a month for Google AI Plus, $19.99 for Google AI Pro, and $99.99 to $199.99 for Google AI Ultra. Build with the API instead and you pay per token, starting around $0.10 per million. The catch is that Google reshuffled and repriced these plans more than once in 2026, so a lot of the Gemini pricing guides you will find quote numbers that no longer exist. Below is every current price, reconciled from each vendor's own pricing page, plus the question the numbers exist to answer: which plan, if any, you should actually pay for.
![Gemini's four consumer plans side by side, from free to Google AI Ultra, at July 2026 US prices.](/blog/gemini-pricing/gemini-plan-prices.png)
The four consumer tiers at US prices as of July 2026 — the API is a separate, per-token price book.
## How Much Does Gemini Cost? Every Plan at a Glance Here is the whole consumer lineup in one place, at US prices as of July 2026.
PlanPrice (US, per month)StorageKey model accessBest for
Free$015 GBGemini 3.6 Flash, limited daily ProCasual questions and everyday use
Google AI Plus$4.99400 GBHigher limits, Gemini Omni and FlashLight users who want more headroom
Google AI Pro$19.995 TBGemini 3.1 Pro, full Deep ResearchDaily and power users (the mainstream paid plan)
Google AI Ultra$99.99 to $199.9920 TB and upHighest limits, Deep Think, agent featuresHeavy creators and developers
Two things cause most of the confusion. Gemini is sold two completely different ways. One is the consumer app, where you pay a flat monthly subscription, which is the table above. The other is the developer API, where your software pays per token for whatever it sends and receives, with no subscription at all. Most people reading a pricing page want the first one. If you are wiring Gemini into your own product, skip down to the API pricing. There is also a third path for organizations: Gemini is sold per seat inside Google Workspace business and enterprise plans, which is a separate price book from the consumer tiers here. The second point: the paid plans are bundled into Google One, the same subscription that sells cloud storage. That is why each tier comes with a big storage allowance attached, and it is also why canceling trips people up. We come back to that at the end. All the consumer plans bill monthly, with no meaningful annual-prepay discount, so the monthly figure is the figure. For the bigger picture of what these models actually are, our guide to [what Google Gemini is](https://geotoolbox.ai/blog/what-is-gemini) covers the lineup in plain terms. ## Is Gemini Free? What the Free Tier Actually Gives You Yes. Gemini has a genuinely useful free tier, and for most people it is enough. On the free plan you get the Gemini app at no cost, with Gemini 3.6 Flash as the default model and a daily allowance of the stronger Pro model for harder questions. You also get image generation, voice mode, 15 GB of shared Google storage, and a small number of Deep Research reports. The honest catch is throttling. Heavy use of the Pro model hits daily caps, and during busy periods the app can quietly drop you from Pro back to Flash. What "limited Pro access" actually means is the part most pricing guides skip, and it is the number that decides whether you need to pay. Google does not publish an exact figure, and the caps shift, but the pattern users report is a handful of Pro prompts and one or two Deep Research reports a day before you are bumped back to Flash. Ask a few questions a day and you will never hit it. Run document-heavy research all day and you will hit it fast, which is the real signal to consider a paid plan. The other free-tier ceiling is the one almost every comparison gets wrong: the context window is gated by plan, not shared across them. Per [Google's own limits page](https://support.google.com/gemini/answer/16275805), the Gemini app gives you 32K tokens with no AI plan, 128K on Google AI Plus, and 1 million on AI Pro and AI Ultra. Google's own yardstick puts 1 million tokens at about 1,500 pages of text; the free tier's 32K is roughly 50. So the headline million-token figure you see quoted for Gemini is a paid feature, and if your reason for considering a paid plan is dropping whole documents into one prompt, that is the number that decides it, not the daily Pro allowance. There is also a free tier on the developer side. The Gemini API gives you free, rate-limited access to most models through Google AI Studio with no credit card. The trade is that free-tier traffic can be used to improve Google's products, so it is fine for prototyping and wrong for anything confidential. Paid API usage, by contrast, is not used to train Google's models. More on that in the API section below. ## Google AI Plus ($4.99 a Month): The Cheap Entry Tier Google AI Plus is the budget paid plan, and it just got cheaper. In June 2026, Google [cut Plus from $7.99 to $4.99 a month and doubled its storage from 200 GB to 400 GB](https://9to5google.com/2026/06/08/google-ai-plus-price-drop/), part of a broader [price war among AI subscriptions](https://techcrunch.com/2026/06/09/google-just-fired-a-warning-shot-in-the-ai-subscription-price-wars/). For $4.99 you get roughly double the free-tier usage limits, a 128K context window (up from 32K), access to the [Gemini Omni](https://geotoolbox.ai/blog/gemini-omni) and Flash models, entry-level Deep Research, Veo 3.1 Fast for short video clips, and the 400 GB of storage. It stops short of the top Pro models, which the next tier adds. Who it is for: light users who keep bumping into the free-tier caps but cannot justify $20 a month. At $4.99 it is one of the cheapest serious AI subscriptions from a major lab, and the storage alone makes it competitive with a plain Google One plan. ## Google AI Pro ($19.99 a Month): The Mainstream Plan If you have heard of "Gemini Advanced," this is its descendant, though the naming took a detour that trips people up at checkout. Gemini Advanced became Google One AI Premium, and Google's own plan page now says the AI Premium plan "has a new name: Google AI Plus" — the $4.99 tier above, with 400 GB of storage. The full-strength tier that most people mean by "Gemini Advanced" is this one, Google AI Pro at $19.99. So if you are following an old guide, check the storage figure and the price rather than the name: 400 GB means you are looking at Plus, 5 TB means Pro. For that you get access to Gemini 3.1 Pro, the full one-million-token context window, full Deep Research, video generation with Veo 3.1, 1,000 Google Flow credits for Google's creative tools, and roughly four times the free-tier usage limits. Storage jumps to 5 TB, up from the 2 TB it carried earlier in the year at the same price. The plan also brings Gemini into Gmail and Docs, raises the limits in Google's coding tools, and adds the Jules coding agent, an upgraded Gemini Notebook (formerly NotebookLM), and Antigravity, Google's agentic development environment. The number that matters for comparison: $19.99 lands a cent under ChatGPT Plus at $20, so the two are a direct fight. Which one wins depends on whether you live in Google's apps or OpenAI's, and we break that down in [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt). One thing to stop counting on: Google's free-year student offer has closed. Its US student page states the previous offer ended on March 11, 2026 and is no longer available; Google scopes that wording by region, so the exact end date you see may differ outside the US. Either way, guides still promising a free year of Pro for a .edu address are describing a deal that is no longer on offer. Who it is for: daily users who hit the free caps, anyone who wants Deep Research and Veo without the throttling, and people already paying for 2 TB of Google One who can fold storage and AI into one bill. ## Google AI Ultra ($99.99 to $199.99 a Month): The Power Tier Ultra is where the pricing gets genuinely confusing, because there are now two of them. At Google I/O in May 2026, Google [restructured the top of the lineup](https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions/). It introduced a new $99.99 Ultra tier aimed at developers and serious creators, and it cut the existing top tier from $249.99 to $199.99. So the two Ultra prices are not a typo. They are two different plans. The $99.99 tier gives you about five times the Pro plan's usage limits, priority access to Antigravity, 20 TB of storage, a YouTube Premium plan, and early access to Gemini Spark, the 24/7 agent that takes action on your behalf (currently limited to Google AI Ultra subscribers in select countries, and rolling out beyond the US). The $199.99 tier pushes usage limits to roughly twenty times Pro, adds Gemini's hardest reasoning mode (Deep Think), Veo 3.1 Standard for higher-fidelity video, up to 25,000 Google Flow credits, and access to Project Genie. That earlier $249.99 price is the one that made consumers wince, and the cut plus the new mid-tier is Google's answer to it. Even so, Ultra is a professional tool, and almost nobody needs the $200 plan for everyday use. It is built for people generating video at volume or running heavy agentic workflows. If you are not sure you need it, you do not. Who it is for: high-volume creators, developers leaning on Antigravity and the agent tooling, and teams that will genuinely use Deep Think and tens of thousands of credits a month. ## Gemini API Pricing: Pay-Per-Token by Model If you are building software on Gemini rather than chatting with it, you do not buy a plan at all. You pay per [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), the chunks of text the model reads and writes, billed separately for input and output. You can reach the same models two ways: the Gemini Developer API, which is the simplest path, or Vertex AI on Google Cloud, which carries the same per-token rates but adds enterprise billing, data governance, and SLAs. Here are the current pay-as-you-go rates for the models worth using as of July 2026.
ModelInput (per 1M tokens)Output (per 1M tokens)Notes
Gemini 2.5 Flash-Lite$0.10$0.40Cheapest model overall
Gemini 3.1 Flash-Lite$0.25$1.50Prior-generation budget tier
Gemini 2.5 Flash$0.30$2.50Fast, multimodal workhorse
Gemini 3.5 Flash-Lite$0.30$2.50New budget tier (July 2026)
Gemini 3 Flash$0.50$3.00Prior-generation Flash
Gemini 3.6 Flash$1.50$7.50New default workhorse (July 2026)
Gemini 3.5 Flash$1.50$9.00Superseded by 3.6 Flash
Gemini 2.5 Pro$1.25 / $2.50$10.00 / $15.00Higher rate for prompts over 200K tokens
Gemini 3.1 Pro (Preview)$2.00 / $4.00$12.00 / $18.00Top Pro model in the public API, paid only; higher rate over 200K tokens
A few things to read into that table. Flash-Lite is the floor and Pro is the ceiling, and the gap is wide: Gemini 3.1 Pro costs eight times more per input token than 3.1 Flash-Lite, and the older 2.5 Flash-Lite is cheaper still. For classification, extraction, and routine summarizing, the cheap models are usually fine, and routing simple work to Flash-Lite is where most teams find their savings. Output costs far more than input, usually four to eight times, so long-winded responses are where bills balloon. The Pro models also carry a context cliff: once a single prompt crosses 200,000 tokens, Gemini 3.1 Pro input roughly doubles to $4.00 and output climbs to $18.00 per million. Retrieval pipelines that stuff large documents into every call can quietly cross that line on every request. One model to watch: [Gemini 3.5 Pro](https://geotoolbox.ai/blog/gemini-3-5-pro) still has not shipped. It has missed several targeted launch dates after Google reportedly rebuilt the base model, and Google has confirmed no release date, no context window, and no pricing for it. It does not appear on the Vertex AI pricing page (renamed the Gemini Enterprise Agent Platform in April 2026) or in the Gemini API model list. Until it lands, 3.1 Pro Preview is the top Pro model in the general API. Two more changes worth flagging if you are pricing this from an older guide: Google **shut down Gemini 2.0 Flash and 2.0 Flash-Lite on June 1, 2026**, so any table still listing them as live is out of date, and it pulled the Pro models out of the free API tier in April, so experimenting with 3.1 Pro now costs money from the first call. Most current models share a one-million-token [context window](https://geotoolbox.ai/glossary/context-window), so the differences are about speed, reasoning quality, and price, not how much they can read. Google's [developer API pricing page](https://ai.google.dev/gemini-api/docs/pricing) is the source of record and was last updated in early August 2026, so check it before you commit a budget. For the developer's full cost reference, the Standard, Batch, Flex, and Priority service tiers, the free-tier and rate-limit ladder, and a worked example of a real API bill, see our [Gemini API pricing](https://geotoolbox.ai/blog/gemini-api-pricing) guide. ## The Hidden API Costs: Batch, Caching, Reasoning, and Media The per-token rates are only half the API bill. Four things move the real number, and only the first two move it down. **Batch mode cuts every paid request in half.** If a job does not need an instant answer, like overnight enrichment or bulk generation, sending it through the Batch API gets a flat 50% discount in exchange for up to 24 hours of latency. It is the single biggest saving most teams skip. **Context caching pays off when you reuse the same context.** If every call ships the same long system prompt, policy document, or codebase, caching it means you pay roughly 10% of the input rate on the repeats, plus a small hourly storage charge. For chatbots and document tools that reread the same material, the savings are large. **Reasoning tokens are the one that surprises people.** Gemini's thinking models bill their internal reasoning as output tokens, even when the visible answer is short. A model with a low headline rate can therefore cost more in practice than one with a higher rate but terser reasoning. When you compare API prices, the sticker rate is a starting point, not the bill. Then there are the separate meters. Image and video generation are not priced in text tokens at all. Google's image model, the one nicknamed Nano Banana, runs around $0.039 an image, and Veo 3.1 video generation runs roughly $0.40 to $0.60 per second. Audio is metered too: audio input costs more per token than text, and the Live API voice models bill separately again. Grounding a response in live Google Search is its own line: the Gemini 3 family gets 5,000 grounded prompts a month free and then $14 per thousand, while the 2.5 models run $35 per thousand. A habit worth borrowing from finance teams: track cost per API call, not just total spend. Two identical-looking requests can bill very differently depending on context length and output, and averages hide the few expensive calls driving most of the cost. ## Gemini vs ChatGPT vs Claude: Which Is Cheaper? On the consumer side, Gemini is the price leader. Its $4.99 Plus tier undercuts everyone, and its $19.99 Pro plan matches [ChatGPT Plus](https://geotoolbox.ai/blog/chatgpt-pricing) and [Claude Pro](https://geotoolbox.ai/blog/claude-pricing), which both sit around $20. If you are choosing purely on price, Gemini wins the entry tier outright. On the API the picture is more nuanced, and it is where the cheapest sticker price can mislead.
ProviderPaid consumer planRepresentative API model (input / output per 1M)
Google GeminiAI Plus $4.99 / AI Pro $19.99Gemini 2.5 Flash, $0.30 / $2.50
OpenAI ChatGPTChatGPT Plus about $20GPT-5.5, about $5.00 / $30.00
Anthropic ClaudeClaude Pro about $20Claude Sonnet 5, $2.00 / $10.00 (permanent rate)
Gemini's Flash models are among the cheapest credible options for high-volume work. Gemini 2.5 Flash at $0.30 input and $2.50 output is a fraction of what the frontier models charge: GPT-5.5 runs about $5.00 and $30.00, and Claude Sonnet 5 $2.00 and $10.00, its permanent rate (the previously planned rise to $3.00 input was cancelled). The fair fight is Flash against each rival's own cheap tier, but even there Gemini tends to win on price, and at scale that gap is real money. But remember the reasoning-token tax. All three providers charge for the model's internal thinking on their reasoning tiers, so a head-to-head on base rates can flip once you measure how verbose each model's reasoning actually is on your workload. The honest move is to benchmark your real prompts on two or three models before committing, not to pick the lowest number in a table. For the full feature-by-feature picture rather than just price, see our comparisons of [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt) and [Claude vs Gemini](https://geotoolbox.ai/blog/claude-vs-gemini). ## Which Gemini Plan Should You Actually Pay For? Most people overbuy. Here is the honest decision, from the bottom up. **Stay free** if you ask Gemini a handful of questions a day and rarely run Deep Research. The free tier covers it, and you will know you have outgrown it the day the app starts dropping you to the lighter model mid-task. **Pay $4.99 for Plus** if you keep hitting the free caps but do not need the top Pro models, or if the 400 GB of storage alone is worth it next to a plain Google One plan. **Pay $19.99 for Pro** if you use Gemini daily for real work, want the full Pro models and uncapped Deep Research, or already pay for Google One storage and can consolidate. This is the right answer for most paying users. **Pay for Ultra** only if you generate video at volume, lean on the agent tooling, or will genuinely use Deep Think. The $99.99 tier is the sane entry point; the $199.99 tier is for professionals who can name exactly why they need it. **Use the API instead of a plan** if you are building a product. A subscription buys one person a seat in the app; the API meters whatever your software does, and for anything programmatic that is the only model that makes sense. Two practical traps before you buy. First, because the plans are bundled into Google One, canceling Gemini and canceling your storage are not the same action, and people get billed after they think they quit. Cancel on the platform you subscribed through, whether that is the web, the Play Store, or the App Store, and confirm your Google One storage stays at the level you actually want. And if you start on a free trial, it converts to the full monthly price automatically, so set a reminder before it renews. Second, a paid plan does not buy identical features everywhere. Agent features like Gemini Spark are limited to select countries and to the Ultra tier, so a buyer outside the US can pay the same $19.99 or $199.99 for a thinner feature set. ## Why Gemini's Price Tag Matters for Your Brand Here is the part that outlasts any specific price. Whichever tier people pay for, Gemini is no longer just a chatbot they visit. It is the engine behind the AI answers in Google Search. [AI Overviews and AI Mode](https://geotoolbox.ai/blog/google-ai-overviews-seo) run on Gemini and reach billions of people a month, which means the model is already summarizing your category and naming competitors to searchers who never open the app. So the more useful budget question is not which plan you buy. It is whether Gemini mentions your brand at all when someone asks it what to use. That is the gap we help businesses close, and it does not come on a pricing tier. The numbers above will move again, because Google has repriced Gemini repeatedly in 2026 alone, so confirm the current figure on Google's plans page or developer pricing page before you pay. And once you know what Gemini costs, the next question is what it says about you. Our [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) maps where Gemini, Google AI Overviews, and the other AI engines cite sources your brand is missing from, so you can see which conversations to get into. You can learn the broader method in our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). ## Frequently Asked Questions ### How much does Gemini cost per month? Gemini is free to start. Paid consumer plans are Google AI Plus at $4.99, Google AI Pro at $19.99, and Google AI Ultra at $99.99 or $199.99 a month, as of July 2026. The developer API is separate and billed per token, starting at $0.10 per million. ### Is Gemini free? Yes. The free tier gives you Gemini 3.6 Flash, a daily allowance of the stronger Pro model, image and voice, and 15 GB of storage, with throttling under heavy use. The API also has a free, rate-limited tier, though free-tier traffic can be used to improve Google's products. ### Is Gemini Advanced the same as Google AI Pro? Roughly, but the naming is messier than a straight rename. "Gemini Advanced" became Google One AI Premium. Google's plan page now states that the AI Premium plan "has a new name: Google AI Plus," which is the $4.99 tier with 400 GB of storage, while Google AI Pro is the $19.99 tier with 5 TB. If an older guide quotes "Gemini Advanced pricing" at around $20 a month, it is describing today's Pro tier. Go by the price and the storage allowance rather than the plan name. ### Is Gemini cheaper than ChatGPT or Claude? On consumer plans, yes. AI Plus at $4.99 undercuts the roughly $20 ChatGPT Plus and Claude Pro tiers, and AI Pro matches them at $19.99. On the API, Gemini's Flash models are among the cheapest, but reasoning-token charges mean you should benchmark your own workload before deciding. ### What is the cheapest Gemini API model? Gemini 2.5 Flash-Lite at $0.10 per million input tokens and $0.40 output. The newest budget tier is 3.5 Flash-Lite (July 2026) at $0.30 and $2.50, though the prior 3.1 Flash-Lite is a touch cheaper at $0.25 and $1.50. Batch mode halves either rate. ### How do I cancel Gemini without losing my Google One storage? Cancel the AI plan from the same platform you subscribed through, then check your Google One tier afterward. Because the AI subscription and storage are bundled, canceling one does not automatically change the other, which is why people get billed after they think they quit. ## Sources - Gemini Developer API pricing - Google AI for Developers - `ai.google.dev/gemini-api/docs/pricing` - Google AI plans - Google One - `one.google.com/about/google-ai-plans` - Google AI Pro and Ultra subscriptions - Gemini - `gemini.google/subscriptions` - Everything new in our Google AI subscriptions, fresh from I/O 2026 - blog.google - `blog.google/products-and-platforms/products/google-one/google-ai-subscriptions` - Google AI Plus gets price drop to $4.99 and storage bump - 9to5Google - `9to5google.com/2026/06/08/google-ai-plus-price-drop` - Google just fired a warning shot in the AI subscription price wars - TechCrunch - `techcrunch.com/2026/06/09/google-just-fired-a-warning-shot-in-the-ai-subscription-price-wars` --- ## Grok vs Gemini: Which Is Better? An Honest Comparison (2026) > Grok vs Gemini, compared for August 2026: current models, pricing, real-time data, coding, image and video, and which engine actually ends up citing your brand. - Canonical: https://geotoolbox.ai/blog/grok-vs-gemini - Published: 2026-06-26 · Updated: 2026-08-22 Grok and Gemini are built on opposite bets. Grok, from Elon Musk's xAI, wires itself to the live conversation on X and leans fast, blunt, and lightly filtered. Gemini, from Google DeepMind, is a multimodal model stitched into Search, Workspace, and Android. So "grok vs gemini" rarely has one winner. It has a winner per job. This guide gives you that task-by-task answer, kept current: the real model lineups and prices (the part most comparisons get stale), the difference between Grok's X feed and Gemini's Google grounding, and one angle the other comparisons skip. The two engines pull answers from different places, so being visible to one does not mean being visible to the other. If people find your brand through AI, that last point outranks any benchmark. One caveat up front: these rosters move almost monthly. Everything below is dated. When a launch lands, the names change before the conclusions do.
![Scorecard of where Grok and Gemini each lean, from real-time data to value.](/blog/grok-vs-gemini/grok-vs-gemini-scorecard.png)
A winner per job, not one winner: Grok takes the live and cheap lanes, Gemini the multimodal and Workspace ones.
## Grok vs Gemini at a Glance Pick Grok for real-time information off X, a fast and cheap reasoning model, and a chattier tone. Pick Gemini for multimodal work, long documents, Google Workspace, image and video generation, and better free value. On a routine question, most people could not tell the two apart. Start with the names, because that is where comparisons go wrong first. xAI's current flagship is [**Grok 4.6**](https://geotoolbox.ai/blog/grok-4-6), launched August 12, 2026, with the prior **Grok 4.5**, **Grok 4.3**, and **Grok 4.20** for heavier multi-agent reasoning and a cheap **grok-code-fast-1** for coding. Google runs on **Gemini 3.1 Pro** for hard tasks and the newest **Gemini 3.7 Flash** (August 13, 2026), with **Gemini 3.6 Flash** still the app and free-tier default, and Deep Think reserved for its top tier. If a comparison still pits Grok 4 against Gemini 2.5, it is describing 2025.
What mattersGrok (xAI)Gemini (Google)
Best atReal-time info, fast reasoning, blunt toneMultimodal, long docs, Workspace, value
Real-time dataLive X (Twitter) feed, when search is onGoogle Search grounding
EcosystemX / the xAI appGmail, Docs, Android, Search
Main paid planSuperGrok, $30/moGoogle AI Pro, $19.99/mo
Image / videoGrok Imagine / Imagine Video 1.5Nano Banana 2 / Veo 3.1
Content filtersLooser, tightened in Jan 2026Conservative, brand-safe
Best forNews junkies, traders, X power usersGoogle users, researchers, creators
## The Models Behind Each (as of July 2026) Both companies ship a family of models, not one. Knowing which is which saves you from paying for the wrong tier or trusting a benchmark for a model you cannot access. On the xAI side (now operating as SpaceXAI after its 2026 SpaceX merger), [Grok](https://geotoolbox.ai/blog/what-is-grok) 4.6, released August 12, 2026, is the current general-purpose flagship, with a 500K-token context window priced at $2 / $6 per million tokens. It builds on Grok 4.5 (July 8, 2026), the "Opus-class" model trained on Cursor data, which rolled out region by region and, after an EU AI Act review, became fully available across the EU on July 16, 2026. The prior Grok 4.3 remains available with its roughly 1-million-token context, Grok 4.20 adds a multi-agent reasoning variant for harder problems, grok-code-fast-1 is the cheap, fast option for coding, and Grok 4 Heavy is the most powerful reasoning mode, gated behind the priciest plan. [Grok 5 is still in training](https://geotoolbox.ai/blog/grok-5) and not released, so anything you read about it is forecast, not fact. Grok runs in the xAI app, on grok.com, and inside X. On Google's side, [Gemini](https://geotoolbox.ai/blog/what-is-gemini) 3.1 Pro handles the heavy reasoning, Gemini 3.7 Flash (August 13, 2026) is the newest fast model with 3.6 Flash still the default most people use in the app, and Deep Think is the slow, high-effort mode reserved for the Ultra plan. Context runs to about 1 million tokens. Gemini is everywhere Google is: the Gemini app, Search's AI Mode, Workspace, and Android.
 Grok (xAI)Gemini (Google)
MakerxAI (founded by Elon Musk)Google DeepMind
Current flagshipGrok 4.6 (Aug 2026); Grok 4.5, 4.3, 4.20 prior-genGemini 3.1 Pro; Gemini 3.7 Flash (newest)
Free modelGrok (limited prompts, basic media)Gemini 3.6 Flash
Top reasoning modeGrok 4 HeavyDeep Think (Ultra plan)
Context window500K (Grok 4.5/4.6); ~1M on Grok 4.3About 1M tokens
Lives inxAI app, grok.com, XGemini app, Search, Workspace, Android
A warning that applies to every section below: the benchmark numbers you see online almost always describe an older model than the one you would pay for today. The cleanest public head-to-head is still Gemini 3 Pro versus Grok 4.1, which shipped a day apart in November 2025. Grok 4.3 and Gemini 3.5 arrived later without a tidy side-by-side, so treat the scores as a recent snapshot, not this week's truth. ## Real-Time Information: Grok's X Feed vs Gemini's Google Grounding This is the single biggest reason people pick one over the other, and it is also the most over-sold. Grok's edge is real, but it is not magic. Grok's advantage is a direct line into X. When you ask about a breaking story, a trending post, or live sentiment, Grok can pull recent X posts with timestamps and links, which a search index can lag by hours. That makes it the stronger pick for traders, journalists, and anyone tracking a moment as it happens. The catch is the part most write-ups miss. By default, Grok has no live knowledge at all. Per [xAI's own model docs](https://docs.x.ai/docs/models), "Grok has no knowledge of current events or data beyond what was present in its training data," and you have to enable its Web Search or X Search tools for it to go look. The real-time feed is a feature you switch on, not an always-current brain. The same caveat applies to its reputation for an unfiltered raw feed: Grok surfaces what it judges relevant from X, not a complete chronological timeline, so it can quietly miss context. Gemini takes the other route. It grounds answers in Google Search and, on the consumer app, can read across your Gmail, Docs, and Drive when you let it. For broad current-events questions with cited sources, that breadth tends to beat Grok; for "what are people saying on X right now," it does not. We cover this split in more depth in our [Grok vs ChatGPT comparison](https://geotoolbox.ai/blog/grok-vs-chatgpt), where the same real-time-versus-grounding tradeoff plays out. ## Coding and Reasoning: Who Is Actually Smarter Here is where readers want a clean winner and the data refuses to give one. Gemini holds a slight edge on hard logic and reliable code, while Grok wins on speed, cost, and a couple of softer skills. On the most-cited public clash, Gemini 3 Pro reached 1501 on the LMArena leaderboard in November 2025, the first model past 1500, just ahead of Grok 4.1 Thinking at 1484. The two rarely run the same test suites, but where the numbers overlap the pattern is consistent. A [llm-stats side-by-side](https://llm-stats.com/models/compare/gemini-3-pro-preview-vs-grok-4.1-thinking-2025-11-17) shows Gemini 3 Pro leading on hard science and math, while Grok 4.1 leads on creative writing and emotional intelligence. When [Tom's Guide ran both through nine real prompts](https://www.tomsguide.com/ai/i-just-tested-gemini-3-0-vs-grok-4-1-with-9-prompts-and-theres-a-clear-winner), Gemini took logic, coding, debugging, and nuanced analysis, while Grok took reasoning style, creative writing, factual accuracy, and self-awareness. Gemini won the tally, narrowly. For coding specifically, the pattern repeats: Gemini tends to produce cleaner, more complete code and handles large codebases well, while several users report Grok generating messy or non-functional output on routine tasks. Grok's counter is economics. grok-code-fast-1 is cheap enough to run aggressively for scaffolding, bug fixes, and test generation, and its API flagship undercuts most rivals on price. If you want one model to lean on for serious development, the case still points to Gemini; if you want a fast, cheap pair-programmer for high-volume iteration, Grok earns its keep.
Benchmark (Nov 2025)EdgeDetail
Overall preference (LMArena Elo)Gemini 3 Pro1501 vs Grok 4.1's 1484; first model past 1500
Math (AIME 2025)Gemini 3 ProNear-perfect score
Science (GPQA Diamond)Gemini 3 Pro91.9% on graduate-level questions
Creative writing (v3)Grok 4.1Leads Gemini 3 Pro on the Creative Writing v3 board
Emotional intelligenceGrok 4.1Leads EQ-Bench on roleplay and empathy
Hallucination (self-reported)Grok 4.1xAI reports ~4%, down from ~12%; see the trust caveat below
Read that table as a map of temperaments, not a scoreboard. Gemini is the careful analyst; Grok is the quick, expressive generalist, and both descriptions predate the models you would run today. **Grok 4.5**, launched July 8, 2026, sharpened xAI's standing without topping the field: on the Artificial Analysis Intelligence Index it landed around 13th of 190 models (as of late July 2026), behind frontier leaders like Claude Opus 5, Fable 5, and GPT-5.6 Sol, not ahead of them. The current flagship, **Grok 4.6** (August 12, 2026), climbs far higher, matching GPT-5.6 Sol as the third-best model on that index. The August 2026 [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) is now the current flagship. For the reasoning-heavy end of this debate, our [Claude vs Gemini comparison](https://geotoolbox.ai/blog/claude-vs-gemini) covers a third contender that often beats both on careful analysis. ## Multimodal, Image, and Video Generation Gemini was built multimodal from the start, and it shows. It reads text, images, audio, and video in a single prompt, so you can hand it a recorded meeting, a chart, or a long clip and ask questions about it. Grok was text-and-image for most of its life, but Grok 4.3 added native video input, so the old line that "Grok is blind and deaf to media" is outdated. Gemini's range is still broader, especially for audio and long video, but the gap narrowed. On generation, both engines now make images and video, which kills another stale talking point. The split is about taste versus polish. ### Images For images, Gemini's Nano Banana family (Nano Banana 2 as the default for everyone, Nano Banana Pro as a higher-fidelity option on paid AI Pro/Ultra) leans toward photorealism, anatomical accuracy, character consistency, and up to 4K output. [Grok Imagine](https://geotoolbox.ai/blog/grok-imagine) leans cinematic and stylized, generates faster, and applies fewer restrictions, though it tops out lower on resolution. People who want a believable product shot pick Gemini; people who want dramatic, moody concept art often prefer Grok. One practical difference: Gemini refuses to generate images of real public figures, where Grok is more permissive. ### Video Video is where this got genuinely competitive in mid-2026. Gemini's Veo 3.1 leads on resolution, up to 4K, and on physical realism, and it plugs into Google's wider creative stack. But Grok Imagine Video 1.5, which reached general availability in June 2026, generates longer base clips with native synchronized audio and, by xAI's own account, topped an image-to-video arena leaderboard at launch. So call the video crown contested: Gemini for polished, high-resolution output, Grok for fast, audio-native social clips. For stills with a looser hand on the filters, Grok is still the more permissive tool. ## Pricing and Plans Both restructured their plans through 2026, so old price tables mislead. Here is where they stand now, with the usual warning that these numbers move. Both have a free tier, and for a lot of people it is enough. Grok's free plan gives limited prompts plus basic image and video generation. Gemini's free plan runs on 3.6 Flash with a 32K-token context (the 1-million-token window is a paid AI Pro/Ultra feature). It is still the more generous free option for everyday work, but for model access, not context: Google lets free users reach Flash and Pro on capped volume. The paid plans are where the value gap opens. SuperGrok is $30 a month; Google AI Pro is $19.99 a month and, per [Google's subscription rundown](https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions/), bundles higher limits with storage and other Google perks. That makes Gemini the better-value mainstream plan, the verdict most reviewers land on, unless live X data is core to your work. At the top, Grok's SuperGrok Heavy runs $300 a month for the Heavy reasoning mode, while Google AI Ultra sits in the $100 to $200 range. Our [Gemini pricing guide](https://geotoolbox.ai/blog/gemini-pricing) breaks down every Google tier, and our [Grok pricing guide](https://geotoolbox.ai/blog/grok-pricing) does the same for SuperGrok.
TierGrok (xAI)Gemini (Google)
FreeLimited prompts + basic mediaGemini 3.6 Flash, 32K context
Main paidSuperGrok, $30/moGoogle AI Pro, $19.99/mo
Top tierSuperGrok Heavy, $300/moGoogle AI Ultra, $100-$200/mo
API flagshipGrok 4.6, $2 / $6 per 1M tokensGemini 3.x, usage-based
On the API side, Grok's flagship pricing is aggressive: roughly half what many comparable frontier models charge. For high-volume automation that runs to tens of millions of tokens a month, that adds up, though Gemini's cheaper Flash-class models compete hard at the low end. ## Content Filters, Bias, and How Much to Trust the Answers Grok built its name on being the less-filtered chatbot, and Gemini built its name on being the safe one. Both descriptions are still roughly true, but the gap is smaller than it was, and the reputations cut both ways. Grok tightened sharply in early 2026. After a scandal over non-consensual sexual images of real people, xAI restricted image generation and cracked down on explicit content, and many users complained the update went too far. As [one report on the change put it](https://piunikaweb.com/2026/01/20/grok-censoring-everything-users-say/), Grok started blocking prompts that were not remotely unsafe, and the roleplay and creative-writing crowd that came for an unfiltered tool felt blindsided. Grok is still more permissive than Gemini, but the "anything goes" era is over. Gemini stays the conservative choice, with more guardrails and more refusals. That suits schools, brands, and regulated settings, but the flip side is real: Gemini also refuses plenty of benign prompts, which is a steady source of user frustration. On political bias and trust, tread carefully, because neutrality is hard to measure and the marketing tends to outrun the evidence. xAI pitches Grok as a "truth-seeking" model, but independent bias tests have pulled in different directions on it, while Gemini leans toward cautious both-sides answers. Treat any "this one is neutral" claim, from either side, as a marketing line rather than a settled finding. The same skepticism applies to accuracy. Hallucination studies flatly contradict each other, with one crowning Grok the most reliable and another putting it well behind, and every reasoning model still misfiring on hard facts more often than its makers admit. There is no trustworthy single number here. The only safe habit with either model is to verify anything that matters against a primary source. ### Privacy and Your Data One more axis worth checking before you commit: what each does with your chats. Both default to using consumer conversations to improve their models, and both offer an opt-out buried in settings. Grok's training on public X data has drawn regulatory scrutiny, while Gemini's reach into Gmail and Docs makes its permissions worth reading before you grant them. For anything sensitive, the business and API tiers of each change the retention story, and that is the version to use. ## Ecosystem, Apps, and Everyday Reliability For a lot of people, the decision never reaches benchmarks. It comes down to where you already work. Gemini's ecosystem is its moat. It is wired into Gmail, Docs, Sheets, Calendar, and Android, and Gemini Live can take real actions across them, like pulling a detail from your inbox or drafting in a document. If you live inside Google's tools, that integration is hard to give up, and reviewers regularly admit Grok beat Gemini on specific tests yet stuck with Gemini for exactly this reason. Grok's home turf is X. If your day runs through the feed, having the assistant right there is the mirror-image advantage. Voice and mobile tilt toward Gemini for everyday use. Grok's voice mode is quick and free on iOS but gated behind a paid plan on Android, while Gemini Live is built into the platform and doubles as a hands-free productivity assistant. Two honest caveats keep this from being a clean Gemini win. First, reliability: Gemini users report mid-session resets and the model occasionally degrading as if it rolled back a version, which is a real friction that benchmark wins do not capture. Second, feel: Grok reads as witty and direct, closer to talking with an opinionated friend, while Gemini is often described as smart but robotic and verbose. If rapport matters to how you work, that difference is not cosmetic. Our [Grok vs Claude comparison](https://geotoolbox.ai/blog/grok-vs-claude) digs further into the personality side of model choice. ## What Grok vs Gemini Means for Your Brand's Visibility Most comparisons stop at "which should I use." If you are marketing a business, the more useful question is the reverse: which of these engines will tell its users about you. And here Grok and Gemini behave so differently that being visible in one tells you almost nothing about the other. The reason is the corpus each one draws from. Grok answers from X posts and whatever its web search surfaces, so showing up in the live conversation and on crawlable pages is what gets you mentioned. Gemini answers from Google's search index and, critically, powers AI Overviews and AI Mode, which sit in front of billions of searches. Earning a citation there means being in Google's index and genuinely worth quoting on the question. The layers only partly overlap. Both engines run an open-web search, so a page that ranks in Google is more likely to be discoverable by Grok's web search, though never guaranteed to be surfaced or cited. But the proprietary layers do not transfer: a brand all over X gets no credit inside Gemini, and a brand sitting in Google's index is not automatically in the live X conversation Grok leans on. This is not theoretical for us. Our own pages pull steady traffic from AI engines on terms Google barely ranks us for, which is the whole point of treating [AI search](https://geotoolbox.ai/blog/how-does-ai-search-work) as its own channel. If you want the practical playbook, our guide on [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) covers the work that actually moves citations across engines, not just rankings on one. ## Which Should You Choose? Here is the part most comparisons dodge: most heavy users do not choose. They keep both and reach for whichever fits the task, the same way you would not use a single app for everything. Still, if you need to commit to one, the split is clear. **Choose Grok if you:** - live on X or need real-time, breaking-news and sentiment answers - want a fast, cheap reasoning model for high-volume use - prefer a blunt, conversational tone over a careful one - want a looser hand on content filters **Choose Gemini if you:** - work inside Google Workspace or on Android - need strong multimodal and long-document handling - generate images or video, or want the better free tier and paid value - want the safer, more brand-friendly default If you are torn, the tie-breaker is rarely the model. It is the ecosystem you already use and the kind of answer you need most often. The gap between these two is small enough that ecosystem, price, and feel will decide it for you more than raw scores ever will. ## The Bottom Line Grok and Gemini are not really fighting for the same seat. One is the fast, plugged-into-X model with a looser tone; the other is the multimodal, Google-wired model with the better free value. Re-check in a month, because the version numbers will have changed by then. If you market a brand, there is a more important comparison than the one you just read: whether these engines mention you at all. Since Grok and Gemini cite from different places, you have to check each separately. That is exactly what geotoolbox's [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) does, querying eight engines including Grok, Gemini, and Google AI Overviews to surface the offsite sources each one cites, competitor pages included, and flag where your brand is missing from the answer. Knowing which model is "better" matters far less than knowing whether either one is sending people your way. ## Frequently Asked Questions ### Is Grok or Gemini better in 2026? Neither wins outright. Grok is better for real-time information from X, fast and cheap reasoning, and a blunt tone; Gemini is better for multimodal work, long documents, Google Workspace, and free-tier value. Most power users keep both and switch by task rather than picking one. ### Is Grok or Gemini better for coding? Recent public testing leans Gemini for reliable, clean code and large codebases, and it won the coding prompts in head-to-head reviews. Grok's edge is speed and cost: grok-code-fast-1 is cheap enough to run constantly for routine work. For serious development, Gemini; for fast, high-volume iteration, Grok. ### Which is better for real-time information, Grok or Gemini? Grok, when its X Search and Web Search tools are switched on, because it reads live posts from X that a search index can lag. But it is not always-current by default. Gemini grounds answers in Google Search, which is broader for general current events but weaker for live social sentiment. ### Is SuperGrok ($30) worth it over Gemini's $20 plan? For most people, Gemini's $19.99 Google AI Pro is the better value, with more included and a more generous free tier. SuperGrok at $30 is worth the premium mainly if live X data is central to your work. One caveat: xAI rolls new models like Grok 4.5 out to its tiers in stages and by region (Grok 4.5 was EU-blocked at launch but became fully available across the EU on July 16, 2026), so check which model your plan actually includes before subscribing. If you just want a capable assistant, start with Gemini's free tier. ### Is Grok still uncensored compared to Gemini? Less than it used to be. A January 2026 safety crackdown tightened Grok's content rules sharply, and many users felt it went too far. Grok is still more permissive than the heavily guardrailed Gemini, but the gap has narrowed. ### Which has the bigger context window, Grok or Gemini? Gemini's flagship runs to about 1 million tokens, while Grok 4.6 sits at 500K (the prior Grok 4.3 offers 1M). Both handle long documents and codebases comfortably, but Gemini currently has more headroom. Treat the flagship figure as the practical number. ## Sources - xAI Grok model docs - Grok 4.5/4.3 specs, API pricing, and real-time search tools - `docs.x.ai/docs/models` - llm-stats: Gemini 3 Pro vs Grok 4.1 Thinking - GPQA, AIME, creative writing and EQ-Bench scores - `llm-stats.com/models/compare/gemini-3-pro-preview-vs-grok-4.1-thinking-2025-11-17` - Tom's Guide: Gemini 3.0 vs Grok 4.1, 9 prompts - real-world head-to-head - `tomsguide.com/ai/i-just-tested-gemini-3-0-vs-grok-4-1-with-9-prompts-and-theres-a-clear-winner` - Tom's Guide: Google Gemini 3 explained - Gemini 3 family and Deep Think - `tomsguide.com/ai/google-gemini-3-everything-you-need-to-know` - PiunikaWeb: Grok moderation update backlash - the January 2026 content tightening - `piunikaweb.com/2026/01/20/grok-censoring-everything-users-say` - Google AI subscriptions - Gemini consumer plan structure - `blog.google/products-and-platforms/products/google-one/google-ai-subscriptions` --- ## Gemini Gems: What They Are and How to Build One (2026) > Gemini Gems are reusable custom versions of Google Gemini. What they are, how to build one, what knowledge files really do, and how they compare to custom GPTs. - Canonical: https://geotoolbox.ai/blog/gemini-gems - Published: 2026-06-25 · Updated: 2026-08-09 Gemini Gems are reusable custom versions of Google Gemini that you set up once and reuse, instead of re-typing the same instructions into a fresh chat every time. Think of a Gem as a saved expert: you give it a name, standing instructions, and optionally a few reference files, and it behaves that way every time you open it.
![The three parts of a Gemini Gem: instructions, knowledge files, and a name.](/blog/gemini-gems/gemini-gem-anatomy.png)
Every Gem is three pieces on top of the same Gemini model: instructions, optional knowledge files, and a name.
## What Are Gemini Gems? A Gem is a customized configuration of Gemini: a name, a set of standing instructions, and optional reference files, saved so you can reuse it. Google calls them [custom AI experts](https://gemini.google/overview/gems/), and they are its version of ChatGPT's custom GPTs. You build it once, and from then on that Gem opens with its instructions already loaded, so you skip the setup paragraph you would otherwise paste into every new chat. The honest framing matters here, because the marketing oversells it. A Gem is not a smarter model or a fine-tuned AI trained on your data. It is a saved prompt with a few extras bolted on. That is genuinely useful for anything you do repeatedly, but it will not make Gemini fundamentally better at a task it already struggles with. It removes the re-setup tax, not the model's limits. Under the hood, a Gem runs on whatever [Gemini](https://geotoolbox.ai/blog/what-is-gemini) model your plan gives you, which in 2026 means the Gemini 3 family (a fast Flash model on the free tier, a more capable Pro model on paid plans). There is no separate "Gem model." When Google updates the model on your plan, your Gems use it automatically, because the Gem is a configuration layer sitting on top of the [large language model](https://geotoolbox.ai/glossary/large-language-model), not a model itself. Paid plans simply give you access to more capable models for it to run on. Three pieces make up every Gem: - **Instructions:** the standing prompt that tells the Gem its role, task, and output style. This is where most of the value lives. - **Knowledge files:** optional reference documents the Gem can pull from. Useful, but with real limits we cover below. - **A name:** so you can find and reuse it from your Gems list. ## Are Gemini Gems Free? Yes. Gems are available on the free tier of Gemini, not just paid plans. This is the single most out-of-date claim in older guides. When Gems launched in August 2024 they were limited to Gemini Advanced and Workspace, so a lot of articles still say you need a subscription. That changed: Gems rolled out to free accounts, and [9to5Google reported the free mobile rollout on March 25, 2025](https://9to5google.com/2025/03/25/gemini-gems-free-mobile/). Most eligible free users can now build and use Gems, though availability still varies by age, region, and account type. What your plan changes is not access to Gems, but the model behind them and the usage limits. A free Gem runs on Google's fast everyday model; a paid plan runs the same Gem on more capable models. So if a free Gem feels weaker than you expected, it is usually the model tier, not the Gem itself. If you are weighing an upgrade, our [Gemini pricing guide](https://geotoolbox.ai/blog/gemini-pricing) lays out what each paid plan costs and unlocks.
PlanGems accessModel behind your GemsBest for
FreeYes, create and useThe fast default Gemini modelPersonal use and trying Gems out
Google AI ProYesAccess to the more capable Pro and thinking modelsDaily heavy use, harder reasoning tasks
Business / EnterpriseYes, with admin controlsWorkspace models plus org-level sharing governanceTeams that need to share and govern Gems
One practical note on where you can build them. You can use Gems anywhere the Gemini app runs, including Android and iOS, but you create and edit them on the web at gemini.google.com. The mobile apps are for running your Gems, not building them. ## How to Create a Gemini Gem, Step by Step Building your first Gem takes about two minutes. 1. Open gemini.google.com and click **"Explore Gems"** in the left sidebar to open the Gem manager (Google's [step-by-step Help page](https://support.google.com/gemini/answer/15235603) mirrors this flow). 2. Click **"New Gem"** to start a blank one. 3. Give it a **name** that describes the job, so you can find it later. 4. Write the **instructions**. This is the part that matters most, and we break down how to write good ones in the next section. 5. Add **knowledge files** if the Gem needs to reference specific documents, like a style guide or a product sheet. 6. Optionally pick a **default tool** so every chat with the Gem opens in Deep Research, Canvas, or image mode instead of a plain chat (covered in its own section below). 7. Use the **preview panel** on the right to test it with a real prompt before you commit. 8. Click **"Save"**. Two things trip people up here. First, the preview panel is not the same as saving. You can test a Gem all day in preview, but if you close the tab without clicking **"Save"**, it is gone. If a Gem you "made" has vanished, this is usually why. Second, you do not have to write polished instructions yourself. Google added a rewrite helper (the wand icon) that takes a rough description of what you want and expands it into structured instructions. It is a fast way to get a first draft you can then tighten by hand. After you save, the Gem lives in your Gem manager, where you can edit, pin, or delete it. Edits apply the next time you open the Gem, and each chat with a Gem is its own thread, so changing the instructions does not rewrite a conversation you already started. If you are not sure where to start, copy a premade Gem instead of building from zero. Open one from the Gem manager, make a copy, and edit its instructions. You inherit a working structure and only change what you need. ## Writing Gem Instructions That Actually Work The instructions are where a Gem succeeds or fails. A vague instruction produces a vague Gem, and most disappointing Gems are disappointing because the prompt was thin, not because the feature is weak. A reliable structure is persona, task, context, and format. Tell the Gem who it is, what job it does, what background it should assume, and how the output should look. The more specific each part is, the less you have to correct later. Here is the difference in practice. A weak instruction reads "You are a helpful marketing assistant." A strong one reads "You are a B2B SaaS copywriter. When I paste a feature description, return three landing-page headlines under 10 words, each leading with the customer outcome, no exclamation points, no buzzwords." The second version makes dozens of small decisions for the Gem up front, so you stop re-explaining them. Three habits make instructions noticeably better: - **Write the constraints, not just the goal.** Telling a Gem what not to do (no jargon, no em dashes, never invent statistics) prevents the failure modes you would otherwise fix by hand every time. - **Iterate in the preview panel.** Run three or four real prompts in preview, watch where the output drifts, and tighten the instruction that caused the drift. Treat it like editing, not a one-shot. - **Format the instructions, do not just write them.** The most common complaint about Gems is that the instructions do not stick. A big part of the fix is structure: break a longer instruction block into Markdown headings, bullet lists, and numbered steps instead of one dense paragraph, so the Gem parses each rule cleanly. A fast sanity check that pays for itself is to ask the Gem to restate its instructions back to you, then correct whatever it misread before you rely on it. Keep instructions tight, too. Your instructions and the live conversation share the same [context window](https://geotoolbox.ai/glossary/context-window), the fixed amount of text the model can hold at once, so a bloated instruction block leaves less room for the actual work. ## Knowledge Files: What Works and What Breaks Knowledge files are the most misunderstood part of Gems, and the source of the most common complaint: "my Gem ignores the document I gave it." Understanding why fixes most of the frustration. When you attach a file to a Gem, the Gem does not memorize it. It pulls in the passages that match your question, the same idea as [retrieval-augmented generation](https://geotoolbox.ai/glossary/retrieval-augmented-generation). If your question lines up with what is in the file, you get a grounded answer. If nothing matches, the Gem can fall back to its general training and answer confidently without being grounded in your file. That silent fallback is why a Gem can look like it is ignoring your document, and it is the same gap behind most [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations): no matching source, so the model fills it in. A few practical realities Google does and does not document: - **The limits are 10 files, 100 MB each.** [Google's documentation](https://support.google.com/gemini/answer/14903178) points to up to 10 knowledge files of 100 MB each, so a Gem holds a handful of reference docs, not a library. ChatGPT's custom GPTs allow 20. - **Google Docs and Sheets stay live.** [Docs and Sheets you add from Drive](https://workspaceupdates.googleblog.com/2024/11/upload-google-docs-and-other-file-types-to-gems.html) reflect their latest version, so updating the doc updates what the Gem knows. Other uploaded files are static snapshots by comparison. - **For large document sets, use the right tool.** If you are trying to load dozens or hundreds of documents, a Gem is the wrong container. Google's NotebookLM is built for large source libraries and grounded answers across them, which is a different job from a reusable Gem. ## Give a Gem a Default Tool Beyond the three core pieces, a Gem can also open in a specific tool instead of a plain chat. The Gem editor has a **default tool** dropdown that sets what every new chat with that Gem starts in: no default tool (regular chat), Create image, Canvas, Deep Research, Create music, or Guided Learning. A [step-by-step 2026 walkthrough](https://www.ai-toolbox.co/gemini-management-and-productivity/how-to-use-gemini-gems-create-custom-2026) shows where to find it. Set it once and you stop choosing the mode by hand every time. This is the one extra that changes what the Gem does, not just what it says. A research Gem set to Deep Research opens straight into a planned, cited report; a drafting Gem set to Canvas opens the editing surface. It is the closest a Gem comes to behaving like a workflow rather than a saved prompt. You can automate a repeat task with **Scheduled Actions**, a Gemini feature (launched in 2025) that runs a saved prompt on a daily, weekly, or monthly schedule and returns the result in your Gemini chat, [per Google](https://blog.google/products-and-platforms/products/gemini/scheduled-actions-gemini-app/). It runs a standalone prompt rather than a Gem itself, but you can paste a Gem's instructions into that prompt to get the same recurring output, like a weekly brand-answer check. One caveat: Scheduled Actions is a paid feature (Gemini Pro or Ultra, plus some Workspace plans), capped at around 10 active actions, so unlike Gems themselves it is not part of the free tier. ## Premade Gems Worth Copying Google ships five built-in Gems, and the fastest way to learn the feature is to copy one and read how its instructions are written. The five, [per Google's own blog](https://blog.google/products-and-platforms/products/gemini/google-gems-tips/), are: - **Brainstormer:** generates ideas and concepts when you are stuck - **Writing editor:** reviews tone, clarity, and structure - **Coding partner:** helps write, explain, and debug code - **Career guide:** advice on roles, skills, and professional growth - **Learning coach:** breaks down hard topics into steps Make a copy of whichever is closest to your use case, then narrow it. The Writing editor, for example, becomes far more useful once you paste your actual brand voice rules into its instructions instead of leaving it generic. Google is also experimenting with Gems built from its Opal tool, surfaced as "Gems from Labs," which some coverage calls "Super Gems." These are mini-app style Gems aimed at multi-step workflows rather than a single saved prompt. As of mid-2026 they are experimental and limited in availability, so treat them as a preview, not a stable feature to build a workflow around yet. ## Gemini Gems vs ChatGPT Custom GPTs If you already use ChatGPT's custom GPTs, Gems will feel familiar, with some real gaps in both directions. The short version: Gems are tightly wired into Google's apps, while custom GPTs have a head start on extensibility and distribution. For a fuller side-by-side of the two assistants overall, see our [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt) comparison.
CapabilityGemini GemsChatGPT Custom GPTs
Custom instructionsYesYes
Knowledge filesUp to 10, with live Google Drive syncUp to 20, static uploads
External API actionsNo (connects to Google apps only)Yes, via Actions
Public marketplaceNoYes, the GPT Store
Image generationVia a default-tool setting (open every chat in Create image mode)Yes, built in
SharingYes, Drive-style link and rolesYes, link or store listing
Free tierYesYes, with limits
Gems win on the Google ecosystem. A Gem reaches into Gmail, Docs, Drive, Maps, and YouTube through the account you already use, and you can pull live context mid-chat by typing @ to mention a connected app. For eligible Workspace users, Gems also appear in the side panel of Docs, Gmail, and Sheets, so the custom expert sits right where the work happens. They lose on reach and extensibility, with no public store to distribute a Gem and no Actions-style way to call an outside API. If you want to publish an assistant to strangers or wire it to a third-party service, ChatGPT is ahead. If you live in Google Workspace, Gems are the lower-friction choice. ## Sharing Gems and Fixing Common Problems You can share a Gem with other people. Google [launched Gem sharing on September 18, 2025](https://workspaceupdates.googleblog.com/2025/09/gem-sharing-gemini-app-workspace.html), and it works like sharing a Google Drive file: open the Gem manager, hit share, then add people by email or send a link, with viewer or editor access. Older guides that say Gems cannot be shared are out of date. A few things still block sharing, and they explain most of the "I can't share my Gem" complaints. If the Gem uses knowledge files that are not shareable, the share option is grayed out. In a Workspace organization, an admin can disable sharing or restrict it to your domain, so a Gem that will not share is sometimes an org policy, not a bug. And when you share a Gem with attached files, you are prompted to share those files too, so recipients can actually use it. The other recurring problem is the Gem that will not save, the "sorry, we can't save your Gem" message. In practice this usually comes from one of three things: you are signed into a mix of personal and Workspace accounts in the same browser, so the Gem is trying to save to the wrong place; the instructions tripped a content filter, which overly absolute or sensitive wording can do; or you edited an existing Gem and never clicked the update button to commit the change. Sign into a single account, soften the wording, and confirm you actually saved. ## Build a Gem to Watch How AI Describes Your Brand One of the most useful marketing Gems is not a copywriter but a brand-answer checker. Build a Gem whose standing instructions describe your brand, your main competitors, and the questions your buyers ask AI tools. Run those questions once with no knowledge attached, to see what the model says about you on its own, then attach your real product facts and run them again. The gap between the two answers is exactly where AI is getting your brand wrong. Set that Gem's default tool to Deep Research, and every run comes back as a sourced report you can compare over time. In our experience, this matters because [AI search](https://geotoolbox.ai/blog/how-does-ai-search-work) now answers a growing share of brand questions before anyone reaches your site, and what the model says about you is frequently stale or wrong. A checker Gem gives you a fast, repeatable way to spot those gaps without rebuilding the prompt each time. Closing them is a separate job: an [optimize-for-AI-search playbook](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) improves the source material the engines draw on, so the answer they generate next time lands closer to the truth. ## Frequently Asked Questions ### What are gems in Gemini? Gems are reusable custom versions of Google Gemini. You give one a name, a set of standing instructions, and optionally some reference files, then reuse it instead of re-typing the same setup into a new chat. They are Google's equivalent of ChatGPT's custom GPTs. ### Are Gemini Gems free? Yes. Gems are available on the free tier of Gemini, not only on paid plans. They launched as a paid feature in 2024 but rolled out to free accounts in 2025. Your plan changes the underlying model and usage limits, not whether you can create Gems. ### How many knowledge files can a Gemini Gem use? You can attach up to 10 knowledge files per Gem, each up to 100 MB, per Google's documentation. That makes Gems good for a handful of reference documents but a poor fit for a large library. For dozens or hundreds of documents, Google's NotebookLM is the better tool. ### Can you share a Gemini Gem? Yes, since September 2025. Sharing works like a Google Drive file: add people by email or share a link with viewer or editor access. Sharing can be blocked if the Gem uses non-shareable knowledge files or if a Workspace admin has disabled it for your organization. ### Can a Gemini Gem use tools like Deep Research or Canvas? Yes. Each Gem has a default tool setting, so you can make every chat with that Gem open directly in Deep Research, Canvas, image creation, music, or Guided Learning instead of a plain conversation. A drafting Gem set to Canvas, for example, opens the editing surface every time instead of a standard chat. ### Do Gemini Gems use a different AI model? No. A Gem is just saved configuration, not a separate model, so it runs on whatever Gemini model your plan includes and upgrades automatically when Google ships a newer one. Paid plans let you run a Gem on a more capable model for harder tasks. ## Take It Further Gems are worth the two minutes it takes to build one, especially for any task you repeat weekly. Start by copying a premade Gem, write specific instructions, and remember that knowledge files ground the Gem only when your question matches what is in them. The bigger point for marketers is that a Gem is only as accurate as the facts you give it, and so is every other AI engine answering questions about your brand. Before you can shape what AI says about you, the engines have to be able to find and correctly read your site in the first place. You can check that groundwork by running your domain through geotoolbox's [AI readiness check](https://geotoolbox.ai/tools/ai-readiness). ## Sources - 5 tips on getting started with Gems, your custom AI experts (Google) - `blog.google/products-and-platforms/products/gemini/google-gems-tips` - Gem sharing in the Gemini app and Workspace (Google Workspace Updates) - `workspaceupdates.googleblog.com/2025/09/gem-sharing-gemini-app-workspace.html` - Free Gemini users can now access Gems on Android, iOS (9to5Google) - `9to5google.com/2025/03/25/gemini-gems-free-mobile` - Upload Google Docs and other file types to Gems (Google Workspace Updates) - `workspaceupdates.googleblog.com/2024/11/upload-google-docs-and-other-file-types-to-gems.html` - Build custom experts with Gems (Google Gemini) - `gemini.google/overview/gems` - Tips for creating custom Gems (Gemini Apps Help) - `support.google.com/gemini/answer/15235603` - Upload and analyze files in Gemini Apps (Gemini Apps Help) - `support.google.com/gemini/answer/14903178` - How to use Gemini Gems: build custom AI assistants (2026) - `ai-toolbox.co/gemini-management-and-productivity/how-to-use-gemini-gems-create-custom-2026` - Gemini app launches scheduled actions to help you stay productive (Google) - `blog.google/products-and-platforms/products/gemini/scheduled-actions-gemini-app` --- ## Gemini vs ChatGPT: Which Is Better? Honest Comparison (August 2026) > Gemini vs ChatGPT, honestly compared and current to August 2026: models, pricing, context, coding, privacy, and which engine actually ends up citing your brand. - Canonical: https://geotoolbox.ai/blog/gemini-vs-chatgpt - Published: 2026-06-25 · Updated: 2026-08-18 Ask which is better, Gemini or ChatGPT, and most answers stall on "it depends." That is true, and useless. So here is the honest version, current as of August 2026: the two are genuinely close, both charge about $20 a month for their main plan, and the right pick changes by the job you need done.
![A category-by-category scorecard of ChatGPT versus Gemini for August 2026.](/blog/gemini-vs-chatgpt/gemini-vs-chatgpt-by-category.png)
Both are excellent and lean different ways — pick by the job, not by an overall winner.
This guide gives you the task-by-task verdict, the current models and prices (the part most comparisons get wrong or leave stale), and one angle most comparisons skip: [Google Gemini](https://geotoolbox.ai/blog/what-is-gemini) and ChatGPT cite sources differently, so being visible in one does not mean being visible in the other. That last point matters more than the benchmark scores if people find your brand through AI. One caveat up front: the model rosters move almost monthly. Everything below is dated. When a launch lands, the names change before the conclusions do. ## Gemini vs ChatGPT at a Glance (August 2026) The short answer: pick ChatGPT for writing, day-to-day reasoning, and the bigger third-party ecosystem. Pick Gemini for long documents, Google Workspace integration, video generation, and bundle value. On most everyday questions, you would struggle to tell the two apart. The model names are the first source of confusion, so start here. OpenAI's current flagship is GPT-5.6, [released on July 9, 2026](https://geotoolbox.ai/blog/gpt-5-6) across ChatGPT, Codex, and the API, with Sol (flagship), Terra (balanced), and Luna (fast) variants. It supersedes GPT-5.5, the prior-generation flagship (until July 2026). Google's lineup, per [Google DeepMind](https://deepmind.google/models/gemini/), runs on Gemini 3.1 Pro for heavy tasks and, as of August 13, 2026, Gemini 3.7 Flash as the newest fast model, shipped just three weeks after 3.6 Flash, while [Gemini 3.5 Pro](https://geotoolbox.ai/blog/gemini-3-5-pro) still has not shipped: Google announced it at I/O on May 19, 2026, but as of August 7, 2026 it remains in partner testing, with no API model ID and no date. If a comparison still cites GPT-5.2 or Gemini 3 Pro as current, it is already out of date.
 ChatGPT (OpenAI)Gemini (Google)
MakerOpenAIGoogle DeepMind
Current flagshipGPT-5.6 (Sol / Terra / Luna)Gemini 3.1 Pro (preview); Gemini 3.7 Flash (newest, Aug 13); 3.5 Pro not yet shipped
Free modelGPT-5.6 Luna (rolling out; Instant still GPT-5.5)Gemini 3.6 Flash
Main paid planPlus, $20/moGoogle AI Pro, $19.99/mo
Top tierPro 5x $100 / Pro 20x $200AI Ultra $99.99 (5x) or $199.99 (20x)
Context windowTiered (about 32K-400K in-app); 1M via APIBy plan: 32K free, 128K AI Plus, 1M on AI Pro and Ultra
Image / videoChatGPT Images 2.0 (gpt-image-2) / no in-app video (Sora app retired Apr 2026)Nano Banana 2 (Pro on paid plans) / Veo 3.1
Best atWriting, reasoning, coding, ecosystemLong docs, Workspace, video, value
The rest of this guide works through where those differences actually bite. ## Pricing: What You Actually Get Free and Paid Both companies restructured their plans in 2026, so old price tables mislead. Here is where they stand now. Start with free, because for many people it is enough. ChatGPT's free tier is mid-transition: OpenAI's help centre said on August 7, 2026 that GPT-5.6 Luna, the fastest and cheapest model in the family, is becoming the default for Free and Go as it rolls out, with GPT-5.5 Instant still behind the Instant setting. Message caps are tight, and neither Free nor Go gets GPT-5.6 Sol; Terra reaches them only inside Codex, never in a standard conversation. Gemini's free tier runs on Gemini 3.6 Flash with a 32K context window. If you mostly need a capable assistant for everyday questions, Gemini's free tier is still the more generous of the two, but not for the reason usually given. It is not the context window; it is model access. Google lets free users reach Flash-Lite, Flash and Pro, capping volume and context rather than model quality, while OpenAI's Free and Go plans get no access to GPT-5.6 Sol at any volume. The main paid plans are effectively tied on price. [ChatGPT Plus](https://chatgpt.com/pricing) is $20 a month and includes GPT-5.6, Deep Research, and image generation. Google AI Pro is $19.99 a month and, per [Google's plan page](https://gemini.google/subscriptions/), bundles expanded Gemini access with 5TB of storage, YouTube Premium Lite, and NotebookLM. Dollar for dollar, Gemini AI Pro packs in more, which is why "Gemini is better value" is the one verdict almost every reviewer agrees on. Our [Gemini pricing guide](https://geotoolbox.ai/blog/gemini-pricing) breaks down every tier and the API token costs. At the top, both run a power tier. ChatGPT Pro is $100 or $200 a month depending on usage limits. Google cut its AI Ultra plan from $249.99 to $199.99 and added a new $99.99 entry tier at I/O 2026, [per Google's announcement](https://blog.google/products-and-platforms/products/google-one/google-ai-subscriptions/). Unless you are running heavy professional workloads, you can ignore both top tiers and miss nothing.
TierChatGPT (OpenAI)Gemini (Google)
Free$0 - GPT-5.6 Luna, limited$0 - Gemini 3.6 Flash, 32K context
EntryGo, $8/moAI Plus, $4.99/mo (400GB)
MainPlus, $20/moAI Pro, $19.99/mo (5TB, YouTube Premium Lite)
PowerPro 5x $100 / Pro 20x $200AI Ultra $99.99 (5x) or $199.99 (20x) (from 20TB, Deep Think)
Is $20 a month worth it over free? For most casual users, no. The free tiers cover daily questions, drafting, and light research. Pay when you hit message caps often, need the heaviest models for work, or want the storage and YouTube extras Gemini bundles in. The $100 to $200 top tiers ([ChatGPT Pro](https://geotoolbox.ai/blog/chatgpt-pricing) and Gemini AI Ultra) are aimed at heavy professional and agentic workloads that need the highest reasoning limits. For everyone else, they are not worth it. ## Writing and Tone: Which Sounds More Human? This is the difference people feel first, and it decides more switches than any benchmark. ChatGPT writes like a person. Gemini writes like a competent report. Spend a week in both and you will see it: ChatGPT varies its rhythm, picks up your tone, and reads warm, while Gemini stays precise, structured, and reliably on point, if a little corporate. A common refrain among heavy users is that Gemini feels like a tool and ChatGPT feels like a colleague. That is a strength for Gemini when you want facts without flourish, like a summary, a spec, or a structured brief. It is a weakness when the writing itself is the product. In hands-on testing, [PCMag found](https://www.pcmag.com/comparisons/chatgpt-vs-gemini-which-ai-chatbot-is-actually-smarter) that ChatGPT follows complex creative instructions more reliably and adds small touches Gemini skips, while Gemini sometimes turns a poem into prose. So the "they are basically the same" myth breaks down fastest here. On a one-line factual question, sure. On a 600-word landing page, a brand email, or a short story, the voice gap is obvious, and it usually favors ChatGPT. If you write for a living, that alone may settle it. For a head-to-head against Anthropic's model on tone, see our [Claude vs ChatGPT comparison](https://geotoolbox.ai/blog/claude-vs-chatgpt). ## Accuracy, Hallucinations, and Sources "ChatGPT is just more accurate" is a myth worth retiring. Which one is more accurate depends entirely on the question. For anything recent or checkable, Gemini has a structural edge: it is grounded in Google Search, so on current questions it usually pulls live information and shows its sources. ChatGPT also searches the web, on by default, but it decides per query whether to look things up rather than grounding every answer, which is where it can produce a confident, wrong, well-written answer about something that changed last month. Both models still [hallucinate](https://geotoolbox.ai/blog/ai-hallucinations), inventing quotes from documents or misreading an image, so neither earns blind trust on facts that matter. The practical rule: for live facts, news, and "what does the latest version do," prefer Gemini and click its citations. For reasoning through a problem you can verify yourself, ChatGPT is its equal. Either way, treat both as a fast first draft of the truth, not the truth itself. This source-grounding gap also explains a deeper split in how each engine decides what to cite, which is where the comparison stops being about chat quality and starts being about whether anyone finds your brand. ## Coding: Benchmark Winner vs Day-to-Day The "Gemini is weaker at coding" line needs nuance. On the public benchmarks the race is mixed and task-specific: Gemini 3.1 Pro leads on some coding and reasoning tests while OpenAI's flagship leads on others, and Gemini is strong at catching logical errors in long files. Yet plenty of developers still reach for ChatGPT. Both things are true. The split is correctness and scale versus readability and explanation. Gemini tends to be thorough and literal, and its huge context lets it reason over an entire repository at once. ChatGPT tends to write cleaner, more idiomatic code and explain its reasoning step by step, which is why beginners and people learning a new language often prefer it. Treat the benchmark scores with caution. They are close, contested, drawn from whichever model versions were tested, and they rarely match how a tool feels after a week of real work. The summary: reach for Gemini when you are analyzing a large codebase or chasing the top score on a given test, and for ChatGPT when you want explained, ready-to-paste code or are still learning. ## Context Window and Long Documents
![Four horizontal bars on a shared 0-to-1M-token scale: Gemini's context window by plan (32K free, 128K on AI Plus, 1M on AI Pro and Ultra) against ChatGPT's tiered in-app window.](/blog/gemini-vs-chatgpt/chatgpt-gemini-context-window.png)
Gemini's 1M-token window starts on AI Pro; its free tier gets 32K, and ChatGPT's in-app window is tiered with 1M reserved for the API.
This is where the spec sheets contradict each other, so let us settle it. Gemini's million-token [context window](https://geotoolbox.ai/glossary/context-window) is a paid feature, not a free one. Per [Google's own limits page](https://support.google.com/gemini/answer/16275805), the Gemini app's context window goes by plan: 32K tokens with no AI plan, 128K on Google AI Plus, and 1 million on AI Pro and AI Ultra. For scale, Google reckons 1M tokens at about 1,500 pages of text or 30,000 lines of code; the free tier's 32K is roughly 50 pages. ChatGPT's in-app context, [per OpenAI's plan details](https://chatgpt.com/pricing), is smaller and tiered: roughly 32K tokens on Plus and up to a few hundred thousand on Pro, with the full 1-million figure reserved for the API, not the chat product most people use. So the comparisons that print "ChatGPT 128K vs Gemini 1M" are about right for everyday users, and the ones claiming ChatGPT matches Gemini's million are quoting the API. The difference is real, not marketing. If your work means dropping an entire research paper, a long contract, a year of transcripts, or a big codebase into one prompt, Gemini handles it without you chopping the file into pieces. For shorter chats, the gap is invisible and you will never notice it. ## Images and Video On images, the lead has changed hands and is now close. Gemini's default generator is Nano Banana 2 (Gemini 3.1 Flash Image), which [replaced Nano Banana Pro in the app in February 2026](https://blog.google/innovation-and-ai/technology/ai/nano-banana-2/); Pro survives for AI Pro and Ultra subscribers as a higher-fidelity regenerate option. The Nano Banana line produces more detailed, photorealistic results and its edits hold the original aspect ratio and resolution better, which is why PCMag handed Gemini the image round in a February 2026 test of GPT Image 1.5 against Nano Banana Pro. Both sides have changed models since, so treat that verdict as a direction, not a current head-to-head. ChatGPT's Images 2.0 model (gpt-image-2, which replaced DALL-E in May 2026) counters with stronger instruction-following and noticeably better text rendering, so it wins for diagrams, infographics, and anything with words baked into the picture. If you are asking "is Nano Banana still better," the answer is: for photo-style images and clean edits, usually yes; for layouts with text, ChatGPT often wins. Video tilts the other way, and not by a little. Gemini generates [multimodal](https://geotoolbox.ai/glossary/multimodal-ai) clips with audio through Veo 3.1 and its Flow filmmaking tools, a full consumer video stack. ChatGPT, meanwhile, has no native in-app video generator at all right now: OpenAI retired the Sora app and website on April 26, 2026 and is [shutting down the Sora API on September 24, 2026](https://help.openai.com/en/articles/20001152-what-to-know-about-the-sora-discontinuation), so this is a full withdrawal rather than a change of shop window. For video generation, Gemini is the only real choice of the two. ## Ecosystem: Google Workspace vs ChatGPT's World For a lot of people, this section decides everything, model quality aside. Gemini lives inside Google. It drafts in Gmail, edits in Docs, cleans up Sheets, pulls from Drive, and rides along on Android, so if your day already runs on Google apps, it removes the copy-paste shuffle entirely. That convenience is real, and it is the most common reason people switch to it. ChatGPT plays a wider field. It has the largest third-party ecosystem: Custom GPTs and the GPT Store, the Codex coding tools, the Canvas editor, its Atlas browser, and broad plugin support through the Model Context Protocol. If your work spans many non-Google tools, or you want to build custom assistants and automations, ChatGPT bends to more shapes. Gemini's answer to custom GPTs is [Gemini Gems](https://geotoolbox.ai/blog/gemini-gems), which are simpler but wired straight into Gmail, Docs, and Drive. Put plainly: Gemini wins if you live in Google Workspace, ChatGPT wins if you live everywhere else. That is often a bigger deciding factor than which model writes a slightly better paragraph. ## Privacy and Your Data Neither is a clear winner here, and the answer is uncomfortable: on their consumer plans, both train on your conversations by default. ChatGPT and Gemini each use your chats to improve their models unless you turn that off in settings, and both let you do so. Google does not train on your Gemini activity inside Workspace apps by default, which is a point in its favor for work accounts. The under-discussed part is retention. Even with history off, providers keep copies for a window, and Google has said Gemini conversations can be human-reviewed and kept for up to three years, so the safe habit is simple: do not paste anything truly sensitive into either one. Gemini also drew scrutiny over its deeper Google reach. Reports in late 2025 said Google had switched on Gemini's smart features across Gmail and Chat by default for many US users, prompting a [class-action lawsuit](https://natlawreview.com/article/silent-switch-new-lawsuit-alleges-google-uses-gemini-ai-secretly-read-gmail-chat), while Google says it did not change those settings and does not train on personal Gmail content. On July 7, 2026 a federal judge dismissed that suit for lack of standing rather than on the merits, granting leave to amend, so it is neither settled nor over. The takeaway holds either way: review the data settings on whichever you choose, because the defaults lean toward collection. ## The GEO Angle: How Gemini and ChatGPT Cite Your Brand
![An illustrative grid showing which engine cites your brand across different prompt types.](/blog/gemini-vs-chatgpt/which-engine-cites-you.png)
The GEO question is per-engine and per-prompt: track where you're cited, not which model is "smarter."
Here is the comparison nobody else runs, and the one that matters most if customers find you through AI: Gemini and ChatGPT do not pull their sources from the same place, so being cited by one is a weak predictor of being cited by the other. Gemini is grounded in Google's search index. Analysis from [Yext](https://www.yext.com/blog/how-chatgpt-perplexity-gemini-claude-decide-what-to-cite), which studied 17.2 million citations across the major engines, found Gemini behaves much like traditional search: it favors official brand websites and established, well-structured sources. ChatGPT's mix shifts more by industry; in some verticals it leans on official brand sites as heavily as Gemini does, in others on third-party reviews and publications. Across all engines, Yext found that verified, directly-distributed brand data (the listings a company controls) made up 54.53% of distinct citation sources, so the brands winning AI mentions tend to own the source of truth, not just the prettiest page. It splits further by query. On a comparison search like this one, our own citation data shows Google's AI leaning heavily on YouTube and Reddit alongside editorial roundups, a very different recipe from the brand-owned pages Gemini favors elsewhere. In our experience at geotoolbox auditing AI visibility, a brand can be quoted confidently by ChatGPT and be invisible in Gemini for the same question, simply because each engine trusts a different kind of source. The practical consequence: optimizing for one engine does not carry to the other. You have to know where each one is sourcing your category and earn a place in both. That starts with [measuring your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) per engine and watching how [Google's AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) cite your space over time. ## Which Should You Use? (Decision by Task) Skip the "it depends" and match the tool to the job. This is the verdict, task by task.
If your main job is...UseWhy
Everyday questions and chatEitherGenuinely interchangeable; pick by free-tier or ecosystem
Writing and creative workChatGPTWarmer voice, follows complex instructions better
Day-to-day codingChatGPTCleaner, better-explained code (huge repos: Gemini)
Studying and exam prepEitherChatGPT explains step by step; for whole-textbook context Gemini needs AI Pro, not the free tier
Long documents and researchGemini1M-token context in the app, on AI Pro and above
Google Workspace workGeminiNative in Gmail, Docs, Sheets, Drive
Image generationTieGemini for photos and edits; ChatGPT for text and diagrams
Video generationGeminiVeo 3.1 and Flow; ChatGPT retired its Sora app
Current facts and newsGeminiGoogle Search grounding with citations
Brand AI visibilityBothEach cites different sources; you need both
Most experienced users land in the same place: use both. Research and fact-check in Gemini, then draft in ChatGPT because it reads more human. With both free tiers, that hybrid costs nothing, and even paying for one and using the other free is common. If your work lives in Microsoft 365 rather than Google Workspace, the third option is worth a look in [Microsoft Copilot vs ChatGPT](https://geotoolbox.ai/blog/microsoft-copilot-vs-chatgpt), and we put Microsoft's assistant against Google's directly in [Copilot vs Gemini](https://geotoolbox.ai/blog/copilot-vs-gemini). Should you cancel one and switch? Switch to Gemini if you want long-context work, Google integration, or better video, and switch to ChatGPT if you want stronger writing and a broader tool ecosystem. A fair number of people who jump to Gemini for the value drift back to ChatGPT for brainstorming and tone, so try the free tier for a week before you cancel anything. ## The One Comparison That Outlives the Model Roster The version numbers in this guide will be stale within months. What will not change as fast is the deeper split: these two engines reason similarly but source differently, and that decides whether your brand shows up when someone asks an AI about your category. If people increasingly discover you through Gemini and ChatGPT instead of a blue-link search, the question stops being "which is better" and becomes "which one is recommending me, and which one has never heard of me." [Geotoolbox](https://geotoolbox.ai/features/citation-interceptor) tracks exactly that: which sources ChatGPT, Gemini, and Google AI Overviews cite for your category, and where your brand is missing from them, so you can earn a place in both engines instead of guessing. Start by [seeing where you stand across engines](https://geotoolbox.ai/features/citation-interceptor) before the next model launch resets the board. ## Frequently Asked Questions ### Is Gemini better than ChatGPT in 2026? Neither is better overall; it depends on the task. Gemini wins on long-document handling, Google Workspace integration, video generation, and raw value. ChatGPT wins on writing quality, day-to-day coding, and its third-party tool ecosystem. On everyday questions, they are close to interchangeable. ### Should I switch from ChatGPT to Gemini? Switch if you live in Google Workspace, regularly work with very long documents, or want stronger video generation and bundled storage. Many people who switch for the value drift back to ChatGPT for brainstorming and tone, so test Gemini's free tier for a week before canceling anything. ### Is the free version of Gemini better than free ChatGPT? For most people, yes, though not for the reason usually cited. Both free tiers are tightly capped on context: Gemini's free plan gets 32K tokens and ChatGPT's in-app window is comparable. Gemini's edge is model access, since Google lets free users reach its Pro model while ChatGPT's Free and Go plans do not include GPT-5.6 Sol at all. ### Which is better for coding, Gemini or ChatGPT? It is close, and the benchmarks split by test rather than crowning one winner. Gemini 3.1 Pro handles entire codebases thanks to its large context, while ChatGPT tends to write cleaner, better-explained code, which is why many developers and learners prefer it day to day. Use Gemini for big-repository work and ChatGPT for explained, ready-to-use snippets. ### Do Gemini and ChatGPT train on my data? Both train on your conversations by default, and both let you turn that off in settings. Google does not train on Gemini activity inside Workspace apps by default. Either way, avoid pasting sensitive information, since chats can be retained for a period and sometimes reviewed to improve quality. ### Can I use both Gemini and ChatGPT together? Yes, and many people do, using one for grounded research and the other for the final draft. With both free tiers, running them side by side costs nothing. ## Sources - Google DeepMind - Gemini models - `deepmind.google/models/gemini` - Google - new Google AI subscriptions (Google One) - `blog.google/products-and-platforms/products/google-one/google-ai-subscriptions` - Google One - Google AI plans - `one.google.com/about/google-ai-plans` - OpenAI - ChatGPT plans and pricing - `chatgpt.com/pricing` - PCMag - ChatGPT vs. Gemini, tested - `pcmag.com/comparisons/chatgpt-vs-gemini-which-ai-chatbot-is-actually-smarter` - Yext - how ChatGPT, Perplexity, Gemini, and Claude decide what to cite - `yext.com/blog/how-chatgpt-perplexity-gemini-claude-decide-what-to-cite` - National Law Review - lawsuit over Gemini's Gmail/Chat default settings - `natlawreview.com/article/silent-switch-new-lawsuit-alleges-google-uses-gemini-ai-secretly-read-gmail-chat` --- ## What Is Query Fan-Out? How AI Search Turns One Question Into Many > Query fan-out is how AI search splits one question into many parallel sub-queries. What Google documents, which engines do it, and what it means for citations. - Canonical: https://geotoolbox.ai/blog/query-fan-out - Published: 2026-06-25 · Updated: 2026-07-25 Query fan-out is the technique where an AI search engine takes one question, quietly splits it into many related sub-queries, runs them in parallel, and writes a single answer from everything it pulls back. It is the reason you are no longer competing for one keyword. You are competing across a whole neighborhood of questions the engine generates on its own. The term comes from Google, and the mechanics are better documented than most of the advice written about them. What Google documents and what the industry has inferred are two different things, and most advice blurs the line. This page keeps it sharp. ## What Query Fan-Out Actually Is **Query fan-out** (sometimes written "query fanout") is one query in, many queries out, one answer back. The engine reads your question, expands it into a set of narrower searches, retrieves results for each, and synthesizes them into a single response with a few citations underneath. Google uses the term in its own documentation. The [Search Central guidance](https://developers.google.com/search/docs/appearance/ai-features) states that "both AI Overviews and AI Mode may use a 'query fan-out' technique, issuing multiple related searches across subtopics and data sources, to develop a response." Google's own [I/O 2025 write-up](https://blog.google/products-and-platforms/products/search/google-search-ai-mode-update/) describes it the same way: "Under the hood, AI Mode uses our query fan-out technique, breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf." The classic example makes it concrete. Ask for "a 5-day trip to Japan" and the engine does not run that one search. It might fan out into searches like hotels in Tokyo, weather in Kyoto in your travel month, rail pass prices, day trips from Osaka, and a dozen other angles you never typed. Each runs as its own retrieval. The answer you read is stitched together from all of them. It is a common opening move in AI answers that use live retrieval, and we cover the full pipeline in our breakdown of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work). Where it runs, fan-out is step one: query understanding and expansion, before any page gets retrieved or ranked. This changes the unit of competition. Traditional search matched your page against one query. Fan-out matches passages from your page against a spread of sub-queries, most of which never appear in any keyword tool. Being the best answer to the headline question is no longer enough if a competitor is the best answer to six of the sub-questions around it. ## "Fan-Out" Already Meant Something Else The word is overloaded, and the confusion is real. Search "fan-out" without "query" in front of it and most of what you find has nothing to do with AI search. In software architecture, **fan-out** is a distribution pattern: one event, message, or write is pushed out to many consumers at once. A social feed that copies a new post into thousands of followers' timelines is doing "fan-out on write." In SQL and BI tools, a "fan-out" is what happens when a join repeats parent rows across child rows (Looker even has a `fanout_on` parameter for it). In electronics, fan-out is how many gate inputs a single output can drive. Same word, unrelated worlds. That overlap is not just trivia. When we pulled the broader content landscape for the bare phrase, the most-cited pages skewed toward electronics and databases, not AI search. The engineering senses still own most of the corpus. The AI-search meaning is the newcomer borrowing an established term, which is exactly why a search for it surfaces distributed-systems documentation first and why this section exists. For the rest of this page, "query fan-out" means the AI-search technique and nothing else. ## How Query Fan-Out Works, Step by Step
![Four-step flow of query fan-out: decomposition, parallel retrieval, merging, and cited synthesis.](/blog/query-fan-out/how-query-fan-out-works.png)
Fan-out in four moves: one question becomes many searches, then one cited answer.
Strip away the marketing and fan-out is four moves. The implementations differ across engines, but the shape holds. 1. **Decomposition.** The model reads your question and generates a set of sub-queries: rephrasings, narrower angles, comparisons, and follow-ups you did not type. This is the "fan" part, and it is generative, so the exact sub-queries change from run to run. 2. **Parallel retrieval.** Each sub-query is fired against an index, matching against [chunked passages](https://geotoolbox.ai/blog/content-chunking) of pages rather than whole documents. Each sub-query returns its own candidate pool, so one prompt produces many overlapping result sets. 3. **Source evaluation and merging.** The system pools the candidates and ranks them. A page that surfaces across several sub-query result sets gets a cumulative advantage, the kind of cross-query reinforcement that hybrid-search systems implement with techniques like reciprocal rank fusion. A passage that answers one narrow sub-query well still has to beat everything else retrieved for it. 4. **Synthesis with citation.** The model fills its context window with the surviving passages and writes one grounded answer, attributing a subset as sources. In the heavier modes, AI Mode's deeper dives and ChatGPT's deep research, this is not a single pass but a loop: the system reads the first round of results and fans out again on what it learned. It is the [retrieval-augmented generation](https://geotoolbox.ai/glossary/retrieval-augmented-generation) pattern with an expansion step bolted onto the front, and it is worth reading [what RAG is](https://geotoolbox.ai/blog/what-is-rag) for how grounding works once the passages are chosen. Two parts of the pipeline decide your fate before the model writes a word: whether the engine generated a sub-query your page can answer, and whether your passage won retrieval for it. That matching leans on semantic similarity, computed with [vector embeddings](https://geotoolbox.ai/blog/vector-embeddings) rather than exact keywords, which is why a passage can be retrieved for a sub-query whose words it never contains. In production this rides on top of Google's full ranking stack, not a standalone vector index. The practical reframe: fan-out turns one ranking contest into a dozen smaller ones, run in parallel, on questions you cannot see in advance. Coverage of the surrounding territory beats a single perfectly optimized page. ## What Google Documents, and What the Industry Inferred Most explainers blur a line worth keeping sharp. "Query fan-out" is Google's term for a Google feature. Everything beyond that is observation, not vendor confirmation, and the honest version of this topic says so. What is documented: Google names the technique for [AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) and [AI Mode](https://geotoolbox.ai/blog/what-is-google-ai-mode), both powered by a custom version of [Gemini](https://geotoolbox.ai/blog/what-is-gemini); as of Google's [I/O 2026 update](https://www.searchenginejournal.com/google-adds-ai-agents-to-search-redesigns-search-box-at-i-o/575311/), the default model behind AI Mode is Gemini 3.5 Flash. Note the hedge in Google's own wording: both features "may use" fan-out, not "always do." Several Google patents describe the underlying machinery. The most on-point, ["Generating query variants using a trained generative model"](https://patents.google.com/patent/US11663201B2/en) (US11663201B2, filed 2018, granted 2023), produces query variants at run time and is widely read as feeding query-refinement features like People Also Search For. A newer filing, ["Search with stateful chat"](https://patents.google.com/patent/US20240289407A1/en) (US20240289407A1), describes generating sub-queries from a user's whole session, not just the latest message. The caveat the patent crowd often skips: Google has never branded any single one of these as "the query fan-out patent," and it does not call the underlying process "query fan-out" inside the filings. The formal patent term is "query variant generation." Google also routinely notes that a patent does not mean the invention is live in production. So these are plausibly related techniques, not a confirmed blueprint of AI Mode's internals. Treat anyone who shows you "the patent that powers fan-out" with mild suspicion. What is inferred: every other engine. ChatGPT, Perplexity, [Copilot](https://geotoolbox.ai/blog/what-is-copilot), and Grok all show fan-out-like behavior, decomposing a prompt into several searches, but their makers have mostly not published the term or the architecture. The behavior is observable; the label is borrowed.
EngineFan-out statusWhat's on the record
Google AI ModeDocumentedGoogle uses the term; powered by a custom Gemini
Google AI OverviewsDocumentedGoogle says it "may use" fan-out to build the answer
Google Gemini (app)ObservedGrounds via Google Search when web search runs; not always labeled "fan-out"
ChatGPT (search)ObservedRuns multiple searches per prompt; OpenAI hasn't published the retrieval architecture
PerplexityObservedMulti-query retrieval matches the pattern; not clearly vendor-labeled
Copilot / GrokAssumedSame RAG family; little public documentation
The engines also differ in how eagerly they fan out. Perplexity, built as an answer engine, searches on nearly every query; ChatGPT and Copilot decide case by case whether a question needs the web at all; Google's AI Mode is the most openly documented, and its deep-research mode can loop into hundreds of searches. Controlled cross-engine comparisons are still thin, so read these as directional. The [engine-by-engine breakdown](https://geotoolbox.ai/blog/how-does-ai-search-work) covers those differences in detail. The takeaway is not that fan-out is fake everywhere but Google. It is that the mechanism generalizes while the certainty does not, so optimize for the behavior, not for a specific vendor's unpublished recipe. ## The Ways a Query Fans Out A prompt does not just spawn random rephrasings. The expansion follows recognizable patterns. One useful framework, from [Dan Petrovic's breakdown](https://dejan.ai/blog/googles-query-fan-out-system-a-technical-overview/) of Google's query-variant patent (US11663201B2), groups the expansion into eight types. The labels are a reading of the patent, not an official Google list, but they map the territory well.
Fan-out typeWhat the engine doesExample from "best CRM for small business"
EquivalentRephrases your query without changing intent"top CRM software for small companies"
Follow-upAsks the logical next question"how much does a small-business CRM cost"
GeneralizationZooms out to the broader topic"what is a CRM"
SpecificationZooms in on a narrower facet"CRM with email automation under 10 users"
CanonicalizationMaps to the standard form of the question"best CRM 2026"
EntailmentSurfaces an implied need you didn't state"does Salesforce integrate with Gmail"
ClarificationResolves an ambiguity in the prompt"CRM for B2B vs B2C small business"
TranslationRuns the query in another language to widen the pool"mejor CRM para pequeñas empresas"
You will notice these map onto things SEOs already track: People Also Ask, related searches, and the questions a thorough topic page answers anyway. That is the point. The vocabulary is new, but the underlying move, anticipate the related questions and answer them, is not. What is genuinely new is the scale, the invisibility, and that the engine can loop, searching again based on what the first round returned. For you, the framework is a coverage checklist, not a tactic. If your page answers the headline query but ignores the obvious follow-up, specification, and entailment questions around it, you are retrievable on one sub-query and absent on the rest. ## How Many Sub-Queries Does One Prompt Trigger? Nobody outside Google knows the exact number, and the published estimates vary enough that you should hold them loosely. Google itself states no figure. What exists is third-party measurement. [Seer Interactive's analysis](https://www.seerinteractive.com/insights/gemini-3-query-fan-outs-research) found an average of roughly 10.7 sub-queries per prompt on Gemini 3, up from about 6 on Gemini 2.5, with some prompts reaching 28. That near-doubling across one model generation is the real lesson: the count is model-version-dependent, not a fixed law of fan-out. Google has since made [Gemini 3.5 Flash the default behind AI Mode](https://www.searchenginejournal.com/google-adds-ai-agents-to-search-redesigns-search-box-at-i-o/575311/) (I/O 2026), a newer model than the ones those counts were measured on, so read them as a snapshot of a moving target, not today's fixed number. A separate [study from Nectiv](https://nectivdigital.com/blog/new-research-we-analyzed-60k-google-fan-out-queries) across more than 60,000 fan-out queries found that 59% of prompts triggered 5 to 11 searches and 24% triggered 12 to 19. At the extreme end, [Ahrefs documented](https://ahrefs.com/blog/query-fan-out/) a ChatGPT deep-research task running hundreds of searches for a single request. So the realistic range is "usually several to a couple dozen, occasionally far more," not a clean number anyone can promise. The more useful finding is what those sub-queries look like as keyword targets: they mostly are not. Multiple analyses report that the vast majority of generated fan-out queries have little or no recurring monthly search volume, which is exactly why [Ahrefs notes](https://ahrefs.com/blog/query-fan-out/) standard keyword tools miss them. They are invented for your prompt, in your context, and may not be generated the same way twice. So you cannot build a list of "the fan-out queries" the way you would for traditional search. Cyrus Shepard calls the belief that you can the "fan-out myth": run the same prompt ten times and you can get ten different sets. The durable target is the theme, not the string: find the commonalities and cover them well enough that whatever variant the engine invents, your page is a plausible answer. ## What This Means for Getting Cited Start with what Google says, because it constrains everything else. Its [guidance is blunt](https://developers.google.com/search/docs/appearance/ai-features): "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." No special files, no AI-only markup, no fan-out schema. The same fundamentals that win traditional search, crawlable pages, genuinely useful content, sound structure, are the price of entry. That floor is also your biggest lever. The strongest predictor of getting cited is still ranking well in classic search: in [Cyrus Shepard's fan-out framework](https://signal.zyppy.com/p/fan-out-framework), AirOps data shows a Google #1 result is cited 43.2% of the time by ChatGPT, far more often than pages outside the top 20. Rank first; fan-out multiplies the visibility you already earned. Plenty of practitioners push back, framing fan-out as a game you can play harder, sometimes describing it as a raffle where you buy more tickets by ranking for more sub-queries. There is a real insight in that, but it cohabits with a real risk, so here is the defensible read. Three things actually follow from how fan-out works: **Cover the neighborhood, on one page.** Answer the obvious sub-questions around your main topic (the fan-out types above), in self-contained sections. This is topical coverage, and it is the closest thing to a direct lever fan-out gives you. The [answer-first restructuring playbook](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) walks through doing it on existing pages. **Do not spin up a page per sub-query.** This is the trap. Generating a thin page for every fan-out variant is exactly the [scaled content abuse](https://developers.google.com/search/docs/essentials/spam-policies) Google's spam policies target, and it backfires. One thorough page with well-structured sections beats fifty fragments, because retrieval works at the [passage level](https://geotoolbox.ai/blog/content-chunking) anyway. A focused page can win a sub-query; a thin doorway page wins a penalty. Reasonable people split on whether to consolidate into one deep page or keep a tight hub-and-spoke cluster, but both camps agree the unit is a real, useful page, not a doorway per query. **Write passages that survive being lifted out.** Retrieval pulls a section away from its page and judges it alone, so each one has to make sense on its own and answer a specific question cleanly. That craft, the [per-passage citability work](https://geotoolbox.ai/blog/ai-content-optimization), is what determines whether you win the sub-query retrieval after fan-out puts you in the running. Does any of this actually move citations? A little, and unpredictably. When [Semrush ran its own test](https://www.semrush.com/blog/query-fan-out-experiment/), broadening four articles to cover fan-out queries, citations rose from two to five. That is a 150% gain and also just three more citations; they peaked at nine, then fell, and the brand's share of voice slipped during the same window. Treat fan-out coverage as a sound bet on a noisy outcome, not a lever with a predictable payout. In our experience tracking which pages get cited across engines, the winners are rarely the ones stuffed with sub-query variants. They are the ones where a single well-structured section happens to be the cleanest answer to whatever the engine asked. The fan-out does the multiplying; your job is to be genuinely good on the underlying topic, not to out-guess the expansion. ## Can You Actually See the Fan-Out? Mostly no, and this is where the tools get oversold. A growing set of "fan-out simulators" and "query fan-out generators" promise to show you the sub-queries an engine will run. Read the fine print on what they actually do: they prompt a language model to generate plausible sub-queries using the patterns from Google's patents. That is an educated guess at the expansion, not a readout of the engine's real internal process. Useful guesses, sometimes. But they are reconstructions, and the gap matters because the real fan-out is generative and, by the patents' design, can be stateful and personalized, varying with context like session or location and changing between runs, so even a good simulator models a moving target. Google's AI Mode exposes some of the sources it drew on; it does not hand you the sub-query list, and ChatGPT and Perplexity expose even less. (The studies earlier could count fan-out queries because Gemini's API grounding returns them as metadata, but that is developer telemetry, not what you see inside consumer AI Mode.) So treat any "we reveal the exact fan-out" pitch with the skepticism it earns, and do not optimize against a simulator's output as if it were ground truth. The honest move is to measure the outcome instead of the mechanism. You cannot watch the fan-out, but you can watch whether engines cite you, for which prompts, and against which competitors, then see whether that improves after you broaden a page's coverage. Google Search Console now has a dedicated [generative AI performance report](https://developers.google.com/search/blog/2026/06/gen-ai-performance-reports) for your impressions in AI Overviews and AI Mode, but it stops there: no clicks, and no sub-queries. [Bing Webmaster Tools](https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview) goes a step further and shows the grounding queries behind your citations (grounding queries fact-check an answer rather than expand it, so they are cousins of fan-out queries, not the same thing, but a rare public window). Neither hands you the consumer fan-out list. You can also spot-check by hand, running your priority prompts in each engine and noting who gets cited, but that does not scale, which is why a dedicated [AI visibility tracking](https://geotoolbox.ai/blog/how-to-track-ai-visibility) approach exists. Citations are noisy run to run, so read trends, not single samples. ## Frequently Asked Questions ### Is query fan-out just query expansion with a new name? It shares the lineage, but it is more than classic expansion. Search engines have padded queries with synonyms and related terms for years. Fan-out goes further: it issues many independent searches in parallel, sometimes loops to search again based on what it finds, then synthesizes one answer with a language model. The family resemblance is real, but calling it "just" query expansion erases the agentic, multi-search loop that makes it new. ### Does ChatGPT use query fan-out? It behaves like it does. ChatGPT with search routinely fires several searches for one prompt and writes a combined answer. But "query fan-out" is Google's term for Google's feature, and OpenAI has not published ChatGPT's retrieval architecture or adopted the phrase. So the accurate statement is that ChatGPT shows fan-out-like behavior, not that it is confirmed to run "query fan-out." ### Can I optimize for the specific sub-queries an engine will generate? No, and tools that promise this are guessing. The sub-queries are generated per prompt, shift between runs, and most have no standing search volume, so there is no fixed list to target. What you can do is cover the theme thoroughly enough that whatever variant the engine invents, your page answers it. ### Is this the same as "fan-out" in coding or SQL? No. In software and databases, fan-out means distributing one message or write to many consumers, or a join that multiplies rows. Query fan-out in AI search is unrelated despite the shared word. If a search for the term returns distributed-systems or BigQuery documentation, you have wandered into the other meaning. ### How do I measure my query fan-out coverage? Indirectly, by measuring citations rather than the hidden sub-queries. Track which prompts cite your pages across engines, which competitors win the ones you lose, and whether broadening a page's coverage moves those results over time. Search Console now shows your AI Overview and AI Mode impressions but not the fan-out queries themselves, so for cross-engine coverage a dedicated AI-visibility tool is the practical route. ### Does query fan-out mean SEO is dead? No. Fan-out multiplies the number of sub-queries you can be retrieved for, but each one still runs against an index your page has to be in, ranked by signals that look a lot like classic SEO. Google is explicit that no special optimization is required, only the usual foundations done well. Fan-out raises the ceiling on how visible you can be; it does not remove the floor. ## Where This Leaves You Query fan-out is the strongest argument yet for an old idea: cover a topic completely, in clean, self-contained sections, on pages a crawler can actually read. The engine multiplies one question into many. You cannot see the expansion, predict the exact sub-queries, or game the simulators that claim to reveal them. What you can do is be the page that answers the neighborhood of questions well enough to get retrieved no matter which way the prompt fans out. And then check whether it worked. geotoolbox's [Content Analyzer](https://geotoolbox.ai/features/content-analyzer) grades a page's citability and shows which engines actually cite it, so you can see whether broader coverage is winning more of the fan-out instead of guessing at sub-queries you will never observe. It cannot tell you exactly why an engine picked a competitor over you, but it measures the thing you actually care about. Watch the trend, since the mechanism is built to stay hidden. ## Sources - Google Search Central: AI features and your website - `developers.google.com/search/docs/appearance/ai-features` - Google: AI Mode in Google Search, I/O 2025 update - `blog.google/products-and-platforms/products/search/google-search-ai-mode-update` - Search Engine Journal: Google I/O 2026, Gemini 3.5 Flash becomes the default model in AI Mode - `searchenginejournal.com/google-adds-ai-agents-to-search-redesigns-search-box-at-i-o/575311` - Google Patents: Generating query variants using a trained generative model (US11663201B2) - `patents.google.com/patent/US11663201B2/en` - Google Patents: Search with stateful chat (US20240289407A1) - `patents.google.com/patent/US20240289407A1/en` - Dan Petrovic (dejan.ai): Google's Query Fan-Out System, A Technical Overview - `dejan.ai/blog/googles-query-fan-out-system-a-technical-overview` - Semrush: We Tested Query Fan-Out Optimization - `semrush.com/blog/query-fan-out-experiment` - Cyrus Shepard (Zyppy Signal): Fan-out Framework, 5 Steps to Improve SEO and AI Visibility - `signal.zyppy.com/p/fan-out-framework` - Google Search Central: Generative AI performance reports in Search Console - `developers.google.com/search/blog/2026/06/gen-ai-performance-reports` - Bing Webmaster Tools: Introducing AI Performance (grounding queries) - `blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview` - Seer Interactive: Gemini query fan-out research - `seerinteractive.com/insights/gemini-3-query-fan-outs-research` - Nectiv: What we learned analyzing 60K+ Google fan-out queries - `nectivdigital.com/blog/new-research-we-analyzed-60k-google-fan-out-queries` - Ahrefs: What is query fan-out? - `ahrefs.com/blog/query-fan-out` - Google: Spam policies for Google web search - `developers.google.com/search/docs/essentials/spam-policies` --- ## Scrunch AI Review (2026): Features, Pricing & Alternatives > An honest Scrunch AI review: what the tool does, its 2026 pricing and Sitecore acquisition, who it's worth it for, and the best Scrunch AI alternatives. - Canonical: https://geotoolbox.ai/blog/scrunch-ai-review - Published: 2026-06-25 · Updated: 2026-08-23 Scrunch AI is one of the more talked-about tools in the AI-visibility space, and as of June 2026 it is also a [Sitecore company](https://www.prnewswire.com/news-releases/sitecore-acquires-scrunch-to-help-brands-influence-discovery-and-buying-decisions-in-the-ai-search-era-302790214.html), a fact most reviews still online have not caught up to. This one is current on that, honest about where Scrunch is strong and where it falls short, and clear about when a cheaper alternative does the same job. One disclosure up front: we build geotoolbox, a competing AI-visibility tool. So we will show our work, name real alternatives, and list our own tool plainly rather than crown it. Read the criticisms with that in mind, and verify the numbers against the vendor's own page before you buy anything. A quick note on the name, because the search results are messy: this is Scrunch AI, the generative engine optimization platform at scrunch.com. It is not the older Scrunch influencer-marketing platform, and it has nothing to do with hair scrunching or scrunchies. ## What Is Scrunch AI? Scrunch AI is an AI search visibility platform. It tracks how your brand shows up when people ask AI assistants questions, and it tells you where you are missing, who is getting cited instead, and what to fix. The company calls itself "The AI Customer Experience Platform," which is marketing for a fairly concrete job: making sure the answer engines see your brand the way you want them to. It sits in the category usually labeled [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), or GEO, sometimes answer engine optimization. The premise is that AI assistants now intercept questions that used to start on Google, so a brand can rank perfectly in classic search and still be absent from the answer a buyer actually reads. Scrunch monitors that gap, and as of July 2026 the self-serve plans cover seven engines (ChatGPT, Claude, Gemini, Perplexity, Google AI Mode, AI Overviews, and Meta), with prompt volume and seats as the tier levers. The company is based in Salt Lake City and was founded in 2023, with Chris Andrew as chief executive. The detail most current reviews miss is who owns it now. In June 2026, Sitecore, the enterprise content and digital-experience software vendor, acquired Scrunch. So the tool you are evaluating is no longer an independent startup; it is the AI-search layer of a much larger martech stack. That shapes a lot of what follows, from roadmap to packaging to who it ultimately serves, so it is worth understanding first. ## Who Owns Scrunch AI? Funding, Growth, and the Sitecore Acquisition Scrunch grew fast for a tool in a category that barely existed two years ago. It launched out of beta in March 2025, and by mid-2025 it had raised a [$15 million Series A](https://www.prnewswire.com/news-releases/scrunch-ai-raises-15-million-series-a-to-rebuild-the-internet-for-ai-consumption-302510915.html) led by Decibel, with Mayfield, Homebrew, and a roster of operator angels including TJ Parker of PillPack, Bryant Chou of Webflow, and Clara Shih of Meta AI. That round followed an earlier $4 million seed and brought disclosed funding to about $19 million at the time; Scrunch's own site now cites roughly $26 million in total backing. By its own count the platform reached more than 500 brands, naming customers like Lenovo, Skims, Crunchbase, Penn State, and Runpod. Then the exit came quickly. On June 3, 2026, [Sitecore announced it had acquired Scrunch](https://www.prnewswire.com/news-releases/sitecore-acquires-scrunch-to-help-brands-influence-discovery-and-buying-decisions-in-the-ai-search-era-302790214.html), for a reported $225 million, per Bloomberg; Sitecore did not officially disclose terms. Sitecore is a long-established digital-experience platform, the kind of content-management and personalization software large enterprises already run. Its chief executive, Eric Stine, framed the deal around acting on AI-search data inside an existing stack, and Scrunch CEO Chris Andrew described it as helping companies "meet buyers where they are, moving beyond traditional SEO." For a 2026 buyer, the acquisition is the single most important fact in this review. It signals the AI-visibility category is consolidating into bigger martech suites rather than staying a field of standalone startups. Practically, it raises a question older reviews never had to ask: will Scrunch keep selling as a self-serve product at its current price, or get folded into enterprise Sitecore deals over time? That is unsettled, and the product still runs under the Scrunch name. But if you are signing up today, you are buying into a Sitecore-owned tool, not a scrappy independent one, so weigh roadmap continuity accordingly. ## What Scrunch AI Does: Core Features Strip away the positioning and Scrunch is built around one loop: watch how AI engines answer questions in your category, find where you are losing, and point you at the fix. It does the watching well. Most of the debate, which we get to below, is about how much of the fixing it actually does for you. The platform tracks brand mentions, sentiment, and share of voice across the engines, then groups related questions into what it calls prompt families so you are looking at intent patterns rather than a thousand isolated queries. You can slice all of it by persona, topic, or geography, which is genuinely useful when one buyer segment sees you and another does not. Underneath that sit citation analysis (which sources the engines pull instead of you), competitor benchmarking, and page audits that flag whether AI crawlers can render your content. The feature most reviewers single out as the standout is the AI bot traffic dashboard, which connects to Google Analytics and shows crawler and referral activity from the engines in a clean, readable view. The newest and least proven piece is the Agent Experience Platform, or AXP, a layer that serves a compressed, machine-readable version of your site to AI agents without changing the human-facing pages.
FeatureWhat it doesCaveat
MonitoringTracks mentions, sentiment, and share of voice across seven engines on the self-serve plansBroad coverage; prompt volume and seats are the tier levers
Prompts & personasGroups queries into prompt families; segments by persona, topic, and geoPowerful, but auto-generated prompts skew too branded; expect manual upkeep
Citations & competitorsShows which sources get cited and how you rank against rivalsStrong for diagnosis; you still act on it elsewhere
Page auditsFlags AI crawlability and rendering issues per URLUseful, but fixes need a developer
AI bot traffic (GA4)Surfaces AI crawler and referral activity in a reporting-ready dashboardThe most-praised feature; lacks page-level and prompt-level breakdown
OptimizeBasic content generation and limited page optimization on the self-serve plans; advanced on EnterpriseA real but shallow execution layer; deeper fixes still need your team
Agent Experience Platform (AXP)Serves a machine-readable version of your site to AI agentsThe roadmap bet; still maturing, and parallel content carries SEO risk
## Scrunch AI Pricing in 2026 Scrunch's [pricing page](https://scrunch.com/pricing/) lists two self-serve plans for brands plus a custom Enterprise tier as of July 2026. One caveat before the numbers: on July 16 we observed the page serving two different lineups to different visitors - some checks (including the one this section describes) got the Starter/Growth ladder below, while rendered-browser checks the same evening got a single four-engine Core plan at $250 with AXP gated to Enterprise. The $250 self-serve entry, 7-day trial, and custom Enterprise top end held constant across both; the plan names, engine counts, and inclusions did not, so treat the breakdown below as one of the two lineups in circulation and confirm which one you're served. In the ladder we documented, Starter runs $250 a month billed annually ($300 month-to-month) with 3 user licenses, 350 custom prompts, 1,000 industry prompts, 3 personas, and 5 page audits, with a 7-day free trial and no credit card. Growth runs $417 a month billed annually ($500 month-to-month) with 5 licenses, 700 custom prompts, 2,500 industry prompts, 5 personas, and 10 page audits. Platform coverage spans ChatGPT, Claude, Gemini, Perplexity, Google AI Mode, AI Overviews, and Meta. Enterprise adds SAML/OIDC security, API access and integrations, a dedicated team, and expanded scale, and agencies get their own track via a toggle on the page.
PlanPriceEnginesIncludes
Starter$250/mo annual ($300 monthly)7 (ChatGPT, Claude, Gemini, Perplexity, AI Mode, AIO, Meta)350 custom + 1,000 industry prompts, 3 personas, 5 audits, 3 seats, 7-day trial
Growth$417/mo annual ($500 monthly)7 (same set)700 custom + 2,500 industry prompts, 5 personas, 10 audits, 5 seats
EnterpriseCustom7 + expanded scaleSAML/OIDC, API access and integrations, dedicated GTM team
The read on that pricing has two parts. First, $250 a month is still steep for the category, where several capable trackers start under $100, so the entry plan makes sense when you will actually use the persona segmentation and audit workflow, not just the mention counts. Second, the packaging has churned repeatedly in 2026: Scrunch has cycled through Starter, Growth, and Core labels, and in mid-July both the four-engine Core layout and the seven-engine Starter/Growth ladder were live simultaneously for different visitors. Any review or listing more than a month old likely describes a retired structure, ours included until this update. Pricing in this category moves fast, so confirm it on scrunch.com before you commit, the same way you would when [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) with any tool. ## Scrunch AI Strengths and Weaknesses
![Two columns splitting Scrunch AI's monitoring strengths from its execution gaps.](/blog/scrunch-ai-review/scrunch-strengths-weaknesses.png)
The recurring verdict across independent reviews: strong diagnosis, thin execution.
The fairest way to read Scrunch is that it is very good at one half of the problem. Here is the split. On strengths, the engine coverage is broad (seven engines on the self-serve plans as of July 2026), the prompt and persona system lets you segment visibility data the way a real marketing team thinks, and the GA4 bot-traffic dashboard is the feature people actually rave about because it turns crawler activity into a slide a client will understand. The diagnostic side is strong too: Scrunch does not just say you are absent, it cross-references answers to suggest why, flagging outdated or contradictory information the models are picking up. It carries SOC 2 compliance and offers hands-on onboarding, which matters for enterprise procurement. None of that is in serious dispute. The weaknesses cluster around a single theme, and it is worth stating plainly because it is the most common complaint across nearly every independent review: Scrunch tells you what is wrong but does little to fix it. Its execution layer is thin. The self-serve plans include only basic content generation and limited page optimization, and the deeper automation, including the Agent Experience Platform and advanced content work, is reserved for Enterprise. It does not repair schema, push fixes into your CMS, or run outreach, so as one third-party review put it, the platform shows you where your brand is invisible, and then it mostly stops. Insights become a backlog. Your team works through it by hand. Reviewers also note that Scrunch does not expose prompt-volume data, so you cannot easily tell which queries are worth the effort, and that the auto-generated prompts lean too branded, leaving you to add and maintain the valuable ones manually. Two more concerns recur: Scrunch tracks a modeled set of prompts on a refresh schedule rather than capturing real user sessions, so a competitor-built review argues the trends can lag what buyers actually see; and several G2 users say the polished dashboards are hard to export, leaving them screenshotting charts, though Looker Studio access on the agency and Enterprise tiers eases that. Two things temper those criticisms. The sharpest of them are published by direct competitors like Profound and Analyze, who have an obvious interest in Scrunch looking incomplete, so weigh their framing accordingly. And the execution gap is partly a category trait, not a Scrunch defect; most pure monitoring tools share it. In our experience building geotoolbox, the deeper issue under all of this is reachability. When a brand shows near-zero visibility, the cause is often a blocked crawler or a JavaScript-only render rather than weak content, and a share-of-voice chart alone cannot tell those two apart. Scrunch's page audits do check crawlability, to its credit, but that signal sits separately from the visibility numbers, so a flat result can still mislead you about the cause. It is worth confirming first. ## Is Scrunch AI Worth It? Who Should Use It Whether Scrunch is worth $250 a month comes down to one question: do you have someone to act on what it finds? The tool produces a strong diagnosis, and a diagnosis is only valuable if it reaches a team that can ship the fix. It is a good fit for mid-market and enterprise brands, and for agencies running visibility for multiple clients, where the persona segmentation, competitor benchmarking, and client-ready reporting earn their keep. If you have an in-house content or engineering function that can turn "we are uncited for these ten prompts" into published pages, Scrunch gives that team a sharp, well-organized worklist. The SOC 2 compliance also clears procurement bars that cheaper tools stay silent on, which matters in regulated industries. It is a poor fit for solo marketers and small teams. If you are one or two people tracking a handful of prompts, you can replicate much of the monitoring by running your own questions through ChatGPT and Perplexity and logging who gets named, and the $250 entry price is hard to justify against that. It is also a stretch if you expected an all-in-one that finds the gap and fills it, because the built-in optimization is basic and most of the filling is still on you. And if revenue attribution is what you need, tying AI mentions to pipeline, that is not Scrunch's strength. If you land in the "skip it" camp, or you want the reachability and citation foundation before you pay for monitoring at all, the alternatives below are organized by the job you are actually trying to get done, the same way our full guide to the [best GEO tools](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools) groups the wider market. ## The Best Scrunch AI Alternatives in 2026 There is no single drop-in replacement, because the right alternative depends on whether your problem is price, action, enterprise depth, or simply whether the engines can reach you. Here are the names that come up most, judged on that basis. ### Scrunch AI vs Profound **Profound** is the alternative Scrunch gets measured against most, so it is worth a direct look. Profound is one of the category's best-funded players, having [raised $96 million](https://geotoolbox.ai/blog/profound-pricing) at a billion-dollar valuation, and it does the things Scrunch only does lightly: daily data refreshes, proprietary prompt-volume data so you know which queries matter, and built-in content workflows. It also starts cheaper, at $99 a month against Scrunch's $250, though its depth and its price climb fast at the top. The short version: Profound is the heavier research platform with the prompt-volume edge, while Scrunch leans on persona segmentation, multi-brand workspaces, and the Sitecore tie-in. Our [Profound alternatives](https://geotoolbox.ai/blog/profound-alternatives) guide goes deeper on where each lands. [**Peec AI**](https://geotoolbox.ai/blog/what-is-peec-ai) is the community's affordable-depth favorite, with a starter plan around €85 a month, a clean interface, and unlimited seats, which makes it the tool teams most often name when they leave a pricier platform. **Otterly.ai** is the cheapest serious starting point, from about $29 a month with a free trial, simple by design. **AirOps** sits at the opposite end of the monitoring-versus-action split: it is built to turn visibility findings into produced content at scale, so it goes further on execution than Scrunch's basic built-in optimization. **AthenaHQ** is the closest enterprise-depth alternative for large organizations, with self-serve pricing around $295 a month. **geotoolbox**, which we make, is the direct budget comparison: plans run from [$99 a month](https://geotoolbox.ai/pricing) ($79 billed annually) with a 7-day free trial — roughly a third of Scrunch's $250-to-$300 entry price — covering visibility tracking across up to eight engines plus the check the category usually buries: whether the AI crawlers can reach and render your pages at all. Scrunch's page audits cover part of that reachability job too, but as a separate signal from the visibility numbers; geotoolbox leads with it, and a free [AI crawler check](https://geotoolbox.ai/tools/ai-crawler-checker) flags the blocks that quietly turn the AI crawlers away before you pay anything. The honest caveat, since it is ours: our prompt-volume intelligence does not match a research platform's real-conversation dataset, so if persona-level research depth is the job, weigh Scrunch and Profound first; if tracking plus reachability is the job, the trial makes the case cheaper than this paragraph can. For the wider field, our [rundown of the best AEO tools](https://geotoolbox.ai/blog/best-aeo-tools) compares the standalone platforms, and our [share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) guide covers the metric most of them sell.
ToolBest forEntry price (mid-2026)Honest note
Scrunch AIEnterprise/agency monitoring$250/mo annual (Starter)Strong diagnosis; thin on execution
ProfoundEnterprise research + workflows$99 Starter / $399 GrowthDaily data, prompt volume; enterprise into the thousands
Peec AIAffordable team depth~€85/moClean UI, unlimited seats; community favorite
Otterly.aiCheapest serious start~$29/moSimple by design; thin for power users
AirOpsAction, not dashboardsQuote-basedTurns findings into produced content
AthenaHQEnterprise depth~$295/moClosest big-org alternative to Scrunch
geotoolboxReachability + trackingFrom $99/mo ($79 annual), 7-day free trialAbout a third of Scrunch's entry price; free crawler check, all 8 engines from the Pro tier (disclosure: ours)
## Frequently Asked Questions ### What does Scrunch AI do? Scrunch AI is an AI search visibility platform that tracks how your brand appears in AI assistant answers across ChatGPT, Claude, Gemini, Perplexity, Google AI, and Meta. It monitors mentions, sentiment, citations, and share of voice, audits whether AI crawlers can read your pages, and flags where competitors are getting cited instead of you. ### Who owns Scrunch AI? Sitecore, the enterprise digital-experience software company, acquired Scrunch on June 3, 2026, for a reported $225 million (per Bloomberg; Sitecore did not officially disclose terms). The product still runs under the Scrunch name, but it is now part of a larger martech stack rather than an independent startup, which is worth weighing when you consider its long-term roadmap. ### Who is the CEO of Scrunch AI? Chris Andrew is the chief executive. He co-founded the Salt Lake City company in 2023 with CTO Robert MacCloy, and both previously worked at Hearsay Systems. Before the Sitecore acquisition the company raised a $15 million Series A led by Decibel, part of roughly $26 million in total backing per its own site. ### How much does Scrunch AI cost? Self-serve entry is $250 a month with a 7-day free trial (no credit card), plus a custom Enterprise tier. Beyond that it depends which lineup you're served: in July 2026 Scrunch's page was showing some visitors a Starter ($250/$300) and Growth ($417/$500) ladder with seven-engine coverage, and others a single four-engine Core plan at $250. Confirm the current plan on scrunch.com. ### Is Scrunch AI worth it? It is worth it for mid-market, enterprise, and agency teams that have the content or engineering capacity to act on what it finds, since its value is in the diagnosis, not the fix. It is hard to justify for solo marketers or small teams tracking a few prompts, who can replicate much of the monitoring manually for free. ### What is the best free Scrunch AI alternative? There is no free tool that fully replaces Scrunch, but you can track a small prompt set by hand by running your questions through ChatGPT and Perplexity and logging who gets named. A free trial like geotoolbox's (which we make) or Otterly's covers basic monitoring, and in our case a free check of whether AI crawlers can reach your site at all. ## The Bottom Line Scrunch AI is a capable, broad monitoring platform that does the watching half of AI visibility well, and as a Sitecore company it now has the backing to keep building. The caveat most reviews land on is real: it shows you the gap and leaves the closing to you, at a price that only makes sense if you have a team to do that work. Before you pay for any monitor, though, rule out a failure that a share-of-voice chart can quietly hide: a page the AI engines cannot fetch. A blocked crawler produces the same flat zero as weak content, and no share-of-voice dashboard can tell the two apart. Our free [AI Readiness scan](https://geotoolbox.ai/tools/ai-readiness) checks it in about two minutes, fetching and rendering your pages the way the AI crawlers do. Confirm reachability first, fix anything it flags, then pick the tracker that matches the job you actually need — and if that job is tracking plus reachability without Scrunch's $250 entry price, the [7-day geotoolbox trial](https://geotoolbox.ai/pricing) is the direct test. The cheapest tool is the problem you avoid paying to measure. ## Sources - Scrunch Pricing - Scrunch, 2026 (current Starter, Growth, and Enterprise pricing, prompts, free trial, engines tracked) - `scrunch.com/pricing` - Sitecore Acquires Scrunch - PRNewswire, June 3 2026 (acquisition announcement, executive quotes, 500+ brands) - `prnewswire.com/news-releases/sitecore-acquires-scrunch-to-help-brands-influence-discovery-and-buying-decisions-in-the-ai-search-era-302790214.html` - Scrunch AI Raises $15 Million Series A - PRNewswire, July 2025 (Series A, investors, Salt Lake City HQ, CEO Chris Andrew, AXP) - `prnewswire.com/news-releases/scrunch-ai-raises-15-million-series-a-to-rebuild-the-internet-for-ai-consumption-302510915.html` - Scrunch | The AI Customer Experience Platform - Scrunch, 2026 (positioning, product modules, named customers) - `scrunch.com` - Scrunch AI on Capterra - Capterra, 2026 (third-party reviews, ratings, legacy pricing reference) - `capterra.com/p/10030499/Scrunch-AI` --- ## What Is Peec AI? Features, Pricing, and Alternatives (2026) > What Peec AI does, its 2026 pricing and funding, who it's for, its real limits, and the best Peec AI alternatives, reviewed by a competing tool's makers. - Canonical: https://geotoolbox.ai/blog/what-is-peec-ai - Published: 2026-06-25 · Updated: 2026-08-24 Peec AI is one of the fastest-growing tools in the AI-visibility category, and most of what is written about it is already out of date. The funding figures, the pricing, even the list of AI engines it tracks have all moved since the popular reviews were published. This explainer is current as of July 2026, honest about what Peec AI does well, where it stops, and how it actually stacks up after a year of fast shipping. One disclosure up front: we build geotoolbox, a competing AI-visibility tool. So we will cite real numbers, link the vendor's own pages, and describe where geotoolbox fits plainly rather than crown it. Read the criticisms with that in mind, and confirm any pricing on peec.ai before you buy. A quick note on the name, since people search for it: it is Peec AI, the AI search analytics platform at peec.ai. The company has not published a meaning behind the word, and no public source shows it to be an acronym, so "what does peec mean" has a short answer: it is a brand name, not a term. ## What Is Peec AI? Peec AI is an AI search analytics platform. It tracks how your brand shows up when people ask AI assistants questions, then tells you where you are missing, who is getting cited instead, and how that changes over time. The company describes it as "AI search analytics for marketing teams," which is a concrete job dressed in plain language: measuring brand visibility inside the answers, not just the blue links. It sits in the category usually called [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), or GEO, sometimes answer engine optimization. The premise is that AI assistants now intercept questions that used to start on Google, so a brand can rank well in classic search and still be absent from the answer a buyer actually reads. Peec watches that gap across the major engines: ChatGPT, Perplexity, Google's AI Overviews and AI Mode, Microsoft Copilot, and Gemini, with more models available on higher tiers. The output is a dashboard built around four numbers: how often you appear ([visibility](https://geotoolbox.ai/glossary/ai-visibility)), where you rank against rivals, how positively you are described (sentiment), and which sources the engines cite to build their answers. Teams use it to benchmark share of voice and find the prompts where they are losing ground. ## Who's Behind Peec AI? Founders, Funding, and Growth Peec AI is a Berlin company, founded in early 2025 by Marius Meiners, Daniel Drabo, and Tobias Siwonia out of the Antler startup program. Marius Meiners is the chief executive, and one detail that says something about the company's competitive streak: TechCrunch notes he is a [former esports athlete](https://techcrunch.com/2026/05/23/peec-one-of-berlins-rising-startups-more-than-doubled-annualized-revenue-in-months-to-10m-sources-say/) who once ranked among the top 100 League of Legends players in the world. The growth has been unusually fast for a category that barely existed two years ago. Peec raised a €7 million seed led by 20VC in 2025, then a [$21 million Series A](https://peec.ai/blog/we-raised-21m-series-a-to-help-brands-win-in-ai-search) led by Singular in November 2025, with Antler, Combination VC, identity.vc, and S20 joining. That brought total funding to around $29 million, and the Series A was raised at a valuation above $100 million. At the raise, Peec reported just over $4 million in annual recurring revenue and 1,300-plus brands and agencies onboarded since February 2025. By May 2026 the company said it had [more than doubled its revenue to about $10 million ARR](https://thenextweb.com/news/peec-ai-berlin-10-million-arr-geo-ai-search) in roughly six months, around 16 months after launch, with 2,500-plus customers and a new New York office. None of that tells you whether the product fits your team, but it explains the pace, and why older reviews mislead on the specifics below. ## What Peec AI Does: Core Features Strip away the positioning and Peec runs one loop: fire a set of prompts at the AI engines on a daily schedule, record how your brand and your competitors show up, and surface where you are losing. It does the watching cleanly, and most of the debate, which comes later, is about how much of the fixing it does for you. The core of the product is brand tracking across engines: visibility share, ranking position, and sentiment for the prompts you choose. Around that sit citation analysis, which shows the exact sources an engine pulled and sorts them into types like editorial, corporate, and user-generated content such as Reddit, plus competitor benchmarking for share of voice against named rivals. Peec also suggests prompts automatically from your website content, and supports multi-country and multi-language tracking. A practical plus: every plan includes unlimited user seats, so you are not paying per head. Reporting runs through CSV exports, a Looker Studio connector, an API on higher tiers, and a Model Context Protocol (MCP) integration that pipes visibility data into tools like Claude and n8n. The newest addition is [AI Shopping Analytics](https://www.globenewswire.com/news-release/2026/06/17/3313204/0/en/peec-ai-launches-ai-shopping-analytics-as-product-recommendations-move-inside-chatgpt.html), launched in June 2026, which tracks product-level visibility: which SKUs an assistant recommends, at what price, and where it sends the buyer. On data collection, reviewers report that Peec reads results from the AI tools' interfaces rather than relying only on official APIs, which is worth confirming with Peec directly if the collection method matters for your procurement.
FeatureWhat it tells youCaveat
Visibility & positionHow often you appear and where you rank for chosen promptsTracks from setup forward; no retroactive history
SentimentHow positively the engines describe you, on a 0-100 scaleA score, not the underlying quotes driving it
Sources & citationsWhich pages the engines cite, grouped by source typeStrong for diagnosis; you act on it elsewhere
Competitor benchmarkingYour share of voice against named rivalsUseful, but limited by your prompt allowance
Suggested promptsAuto-generated prompt ideas from your siteGood starting point; the valuable ones need manual curation
AI Shopping AnalyticsProduct-level visibility, price, and placement in AI answersNew (June 2026); most relevant to ecommerce
IntegrationsCSV, Looker Studio, API, and MCP into Claude/n8nAPI and some integrations sit on higher tiers
## Peec AI Pricing in 2026
![Peec AI's July 2026 plan prices with the three-model limit and per-model add-on costs.](/blog/what-is-peec-ai/peec-pricing-model-limit.png)
Peec's self-serve tiers all cap tracking at three models — extra engines are a paid add-on.
Peec's [pricing page](https://peec.ai/pricing) lists three self-serve brand plans plus a custom Enterprise tier, priced regionally. Billed annually, the brand plans are €85 a month for Starter, €205 for Pro, and €425 for Advanced; in US dollars that is about $80, $205, and $420 (month-to-month runs higher, around $95, $245, and $495). Each comes with a 7-day free trial and no credit card. Older reviews quoting an €89 starter or a €199 pro are reading 2025 pricing.
PlanPrice (EUR/mo)PromptsIncludes
Starter€8550Choose 3 models, 1 project, unlimited users, daily tracking
Pro€205150Choose 3 models, 2 projects, unlimited users, daily tracking
Advanced€425350Choose 3 models, 5 projects, multi-country, Looker Studio
EnterpriseCustomCustomAll models, unlimited projects, API access, SSO, dedicated support
The detail that matters most is the model limit. Every self-serve plan lets you track only three models, chosen from ChatGPT, Google AI Mode, AI Overviews, Microsoft Copilot, Perplexity, and Gemini. Tracking a fourth or fifth engine is a paid add-on that scales with your plan: €30 a month per extra model on Starter, €70 on Pro, and €140 on Advanced. Models like Claude, DeepSeek, Qwen, GPT-5 Search, and Mistral are reserved for Enterprise, which can track up to 11 models. So broad coverage on the Pro plan can add a few hundred euros to the €205 headline, which is the most common surprise in user reviews. Agencies get a separate, credit-based track. Essential is €205 a month for 10,000 credits, which Peec frames as roughly 111 daily prompts, then Growth at €425 for 25,000 credits, Scale at €675 for 65,000 credits, and a custom Comprehensive tier. Pricing in this category moves fast, so confirm the current figure on peec.ai before you commit, the same way you would when [comparing what any AI visibility tool costs](https://geotoolbox.ai/blog/profound-pricing). ## Peec AI Strengths and Limitations The fairest read on Peec is that it is very good at one half of the problem. The interface is clean, setup takes a few minutes, and reviewers praise how quickly a non-technical marketer can stand up a useful dashboard. The multi-engine coverage, the competitor benchmarking, the source-level citation detail, and the unlimited seats are real advantages, and support gets named often as responsive. As a Berlin company it may also be an easier fit for some EU buyers, though you should confirm hosting, subprocessors, and data residency in its own documentation. Temper that praise, though: Peec only launched in 2025, so its public review base is still small. Read the enthusiasm as early signal, not a long track record. The weaknesses cluster around one theme, and it is the most common complaint across nearly every independent review: Peec tells you what is happening but does little to fix it. There is no audit that hands you the specific changes to make, no content generation, and no way to push fixes to your site. Insights become a backlog your team works through by hand. Three more limits recur. ROI attribution is thin, with no native way to tie an AI mention to a click, lead, or sale, though you can rig a rough visibility-to-revenue view through its API and MCP integration. There is no retroactive data, so a new account cannot reconstruct where a brand stood before tracking began. And full engine coverage costs extra, as the pricing section shows. A final caveat applies to the whole category, not just Peec. Large language models are non-deterministic, so the same prompt can return different answers minutes apart. A single daily snapshot can wobble for reasons that have nothing to do with your marketing, which is why the sound way to [measure AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) is to read the trend over weeks, not any one day's score. In our experience building geotoolbox, the problem underneath all of this is that a visibility score tells you the symptom, not the cause. When a brand reads near zero, one cause worth ruling out is reachability: a blocked crawler or a JavaScript-only page can suppress the score even when the content is strong, and a share-of-voice number alone cannot separate that from genuinely weak content. ## Where Peec AI Fits, and Why Reachability Comes First Peec answers a sharp question well: are you showing up in AI answers, and how do you compare to rivals? For a marketing team that already produces content and wants evidence of whether AI search is working, that is genuinely useful. It works best once you have confirmed an earlier step: that the pages you want cited can actually be fetched and read by the AI crawlers. Many sites fail that step without knowing it. They block AI crawlers in robots.txt by accident, serve content that only renders in JavaScript the bots do not execute, or sit behind bot protection that turns those crawlers away. When that happens, a visibility score can read low for a reason that has nothing to do with your content. To its credit, this is a gap Peec closed in 2026. Its [Agent Analytics](https://peec.ai/changelog) module added a free Crawlability checker that tests a domain's robots.txt against more than 40 AI bots, plus Crawl Insights, which reads your server logs to show which AI crawlers actually hit which pages. So reachability is no longer something Peec ignores, and that is worth knowing if you are weighing tools on this feature. The broader point still holds: visibility and reachability are different layers, and reachability is worth confirming first, because no tracker can lift a score for a page an engine never sees. You can run that check inside Peec, or with a standalone free tool such as our [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker), which flags the [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) your site blocks in about two minutes with no account. Once reachability is clean, the work that moves a score is the same whichever tracker you use: read which source types get cited for your prompts, make your key pages reachable and well-structured (whether an [llms.txt file](https://geotoolbox.ai/blog/llms-txt) helps is its own debate), and target the questions where a competitor is named instead of you. ## Peec AI Alternatives and Competitors There is no single drop-in replacement for Peec, because the right alternative depends on whether your problem is price, engine coverage, or enterprise depth. The category is crowded, so the practical way to choose is by the job you are trying to get done, not by a ranking. ### Peec AI vs Profound Profound is the alternative Peec gets measured against most. It is the best-funded player in the category, with daily data refreshes and proprietary prompt-volume data that Peec lacks, and it starts around $99 a month against Peec's roughly $80 (both billed annually). The trade-off is depth versus simplicity: Profound goes deeper on research and prompt-volume data at a higher entry price, while Peec trades some of that depth for a cleaner interface and unlimited seats. Our [Profound alternatives](https://geotoolbox.ai/blog/profound-alternatives) guide compares them in more detail. Beyond that head-to-head, [Scrunch AI](https://geotoolbox.ai/blog/scrunch-ai-review) is the enterprise-leaning option, now owned by Sitecore, priced from $250 a month. Otterly.ai is the cheapest serious start, from about $29 a month. AthenaHQ is the closest enterprise-depth alternative, around $295 a month. For the wider field, our [rundown of the best GEO tools](https://geotoolbox.ai/blog/best-generative-engine-optimization-tools) compares the standalone platforms. geotoolbox, which we make, is also in this space, and the practical difference from Peec is how you pay for coverage: plans run from [$99 a month](https://geotoolbox.ai/pricing) ($79 billed annually) with a 7-day free trial, engine coverage expands with the tier up to all eight rather than through per-model add-ons (Peec's run €30 to €140 each per month), and every plan pairs the tracking with the reachability check — whether AI crawlers can fetch your pages — that this article keeps flagging as the first thing to rule out. A free crawler check covers that layer with no account. The honest caveat, since it is ours: our prompt intelligence does not match a dedicated research platform's dataset, and Peec's unlimited seats on every plan is a real edge for larger teams — weigh both on the same terms.
ToolBest forEntry price (mid-2026)Honest note
Peec AIAffordable team depth€85/mo (~$80)Clean UI, unlimited seats; monitoring-focused, add-ons for more engines
Scrunch AIEnterprise/agency monitoring$250/moSitecore-owned; strong diagnosis, thin execution
ProfoundResearch + workflows~$99/moDaily data and prompt volume; price climbs at the top
Otterly.aiCheapest serious start~$29/moSimple by design; thin for power users
AthenaHQEnterprise depth~$295/moClosest big-org alternative
geotoolboxReachability + trackingFrom $99/mo ($79 annual), 7-day free trialEngines expand by tier (no per-model add-ons), reachability check on every plan (disclosure: ours)
## Frequently Asked Questions ### What does Peec AI do? Peec AI is an AI search analytics platform that tracks how your brand appears in AI assistant answers across ChatGPT, Perplexity, Google AI Overviews and AI Mode, Microsoft Copilot, and Gemini. It measures visibility, ranking position, sentiment, and which sources the engines cite, and it benchmarks all of that against your competitors so you can see where you are winning or missing. ### How much does Peec AI cost? As of mid-2026, Peec's brand plans are €85 a month for Starter, €205 for Pro, and €425 for Advanced, plus a custom Enterprise tier, with a 7-day free trial and no credit card. Each self-serve plan covers only three AI models; tracking more is a paid add-on from €30 to €140 per model per month, so full coverage costs more than the headline price. Agencies have a separate credit-based track starting at €205 a month. ### Who is the CEO of Peec AI? Marius Meiners is the co-founder and chief executive. He started the Berlin company in early 2025 with Daniel Drabo and Tobias Siwonia, and TechCrunch notes he is a former esports athlete who once ranked among the top 100 League of Legends players globally. ### What does "peec" mean? Peec is a brand name, not an acronym. The company has not published a meaning behind the word, so if you are searching for a definition, there is no hidden one to find. It refers to Peec AI, the AI search analytics platform at peec.ai. ### Is Peec AI worth it? It is worth it for established marketing teams and agencies that already produce content and have someone to act on the data, since its value is a clean diagnosis rather than a fix. It is harder to justify for brand-new sites with little content or solo marketers tracking a few prompts, who can replicate basic monitoring by hand, and it does not tie AI mentions to revenue. ### What are the best Peec AI alternatives? The common alternatives are Scrunch AI for enterprise monitoring, Profound for research depth and prompt-volume data, Otterly.ai for a cheap start, AthenaHQ for large organizations, and geotoolbox (disclosure: ours) for tracking with a built-in reachability check from $99 a month. Whichever you weigh, confirm AI crawlers can reach your site first, a check you can run free in Peec's own Crawlability tool or with ours at geotoolbox. ## The Bottom Line Peec AI is a well-built, fast-growing AI search analytics tool that does the watching half of the job well. For an established marketing team with content to defend and someone to act on the data, it is a clean, fairly priced way to see where you stand across the major AI engines. The caveats to weigh are that it monitors rather than fixes and that full engine coverage costs more than the headline price. Whichever tracker you choose, settle reachability first, because an engine cannot cite a page it cannot fetch, and a blocked crawler quietly caps any visibility score. Peec now checks this with its own Crawlability tool, and you can also run a free standalone check with our [AI crawler checker](https://geotoolbox.ai/tools/ai-crawler-checker) in about two minutes. Confirm the engines can reach you, fix what is broken, then let a tracker — Peec, or [geotoolbox's 7-day trial](https://geotoolbox.ai/pricing) if you want reachability and tracking in a single subscription — measure where you stand. ## Sources - We Raised $21M Series A - Peec AI, November 2025 (Series A amount, lead investor Singular, $29M total, $4M ARR, 1,300+ customers, founders) - `peec.ai/blog/we-raised-21m-series-a-to-help-brands-win-in-ai-search` - Peec AI more than doubled annualized revenue to $10M - TechCrunch, May 2026 ($10M ARR, $100M+ valuation, CEO Marius Meiners esports background, Berlin/NY) - `techcrunch.com/2026/05/23/peec-one-of-berlins-rising-startups-more-than-doubled-annualized-revenue-in-months-to-10m-sources-say` - Berlin's Peec AI doubled revenue to $10M ARR - The Next Web, May 2026 (ARR corroboration, Series A, valuation) - `thenextweb.com/news/peec-ai-berlin-10-million-arr-geo-ai-search` - Peec AI launches AI Shopping Analytics - GlobeNewswire, June 2026 (AI Shopping Analytics, 2,500+ customers, founders, funding) - `globenewswire.com/news-release/2026/06/17/3313204/0/en/peec-ai-launches-ai-shopping-analytics-as-product-recommendations-move-inside-chatgpt.html` - Pricing for Peec AI - Peec AI, 2026 (current brand and agency pricing, prompts, models, add-on costs, free trial) - `peec.ai/pricing` - Peec AI Changelog - Peec AI, 2026 (Agent Analytics: Crawlability checker and Crawl Insights server-log AI-crawler tracking) - `peec.ai/changelog` --- ## Agentic Commerce: Will AI Agents Recommend Your Brand? > Agentic commerce in 2026: how AI agents find, judge, and recommend brands, what's live after the ChatGPT Instant Checkout pullback, and how to check yours. - Canonical: https://geotoolbox.ai/blog/agentic-commerce - Published: 2026-06-24 · Updated: 2026-07-20 Agentic commerce is online shopping where an AI agent does the buying for a person: it takes a goal, finds the options, compares them, and completes the purchase. For brands, that rewrites the oldest question in ecommerce. For twenty years the job was to get a human to your page and persuade them. Now the visitor sizing up your product, and increasingly paying for it, is software that does not care about your hero image. The question is no longer whether people can find you, but whether an agent will find, trust, and recommend you, and check out when it does. ## What Is Agentic Commerce? Agentic commerce is a model of online shopping where an [AI agent](https://geotoolbox.ai/glossary/ai-agent) acts as a delegated buyer. The shopper sets the goal and the guardrails ("find a waterproof size-10 hiking boot under $150, delivered by Friday"), and the agent does the searching, comparing, and, when allowed, the buying. You will also see it called a-commerce, and the agents doing the work called AI shopping assistants or shopping agents. Put simply: the agent does not help with the purchase, **it is the buyer**. That is the line from regular ecommerce. A traditional shopper opens ten tabs, reads reviews, and fills in a checkout form. An agent collapses that into one instruction and returns a short list or a finished order. The person may never see your product page or your cart. For a brand, the shift is uncomfortable in a specific way. The thing sizing up your product is a model, not a person, and it weighs your listing on machine-readable signals, not design or copy. The job now is making sure that when an agent shops your category, it can find you, read you correctly, decide you are the answer, and pay you. ## How Big Is Agentic Commerce, Really? Real, but earlier and messier than the headlines suggest. "Agentic commerce" is not defined the same way twice, so the numbers measure different things and land orders of magnitude apart. Read them with the scope attached.
SourceWhat it measuresWhenFigure
Grand View ResearchAgentic commerce market (global)2025 to 2033$5.71B growing to $65.47B (35.7% CAGR)
McKinseyUS B2C retail revenue orchestrated by agentsby 2030up to $1 trillion
McKinseyGlobal goods spend influencedby 2030$3 to $5 trillion
PYMNTSRetailers piloting AI shopping agents202643%
IBMConsumers using AI in the buying journey202645%
The demand signals are more concrete than the sizing. [IBM found](https://www.ibm.com/think/topics/agentic-commerce) 45% of consumers already use AI in the buying journey, Adobe clocked roughly 805% year-over-year growth in AI-referred retail traffic around Black Friday 2025, and Morgan Stanley put Americans who made an AI-assisted purchase last month near 23%. Now the hedge most explainers skip. Interest cooled from a spring-2026 peak, plenty of "agentic" pilots never complete a purchase, and [McKinsey notes](https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-agentic-commerce-opportunity-how-ai-agents-are-ushering-in-a-new-era-for-consumers-and-merchants) adoption tracks "comfort, norms, and credibility" more than raw capability. Treat the trillion-dollar forecasts as a direction, not a budget line, and find out where your brand stands today. ## Can an AI Agent Find Your Brand? Before any of the protocol talk matters, an agent has to answer four questions about you: can it find you, can it understand you, will it trust and recommend you, and can it buy from you. Most coverage jumps straight to checkout. The earlier gates are where brands actually lose. Start with discovery, because nobody has solved it. The payment protocols define how a transaction completes, not how an agent finds your product in the first place. With no agreed discovery standard yet, visibility comes from the unglamorous basics done well. ### Be Reachable The answer-fetching bots behind ChatGPT and Perplexity, and the agents acting on your pages, have to get in and read you. If your prices, specs, or whole product pages only render after JavaScript runs, many of them see an empty page. Our guide to [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) covers which bots matter and how to control them. Before an agent can recommend or buy from you, a crawler has to reach you. The free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows which of the major AI crawlers your robots.txt allows or blocks, with the exact line to fix. ### Be in the Feed Agents lean on structured product data and real-time availability, so a clean, complete product feed (Google Merchant Center, valid GTINs, full attributes) is now table stakes. Often the issue is not a brand problem at all but a catalog-data one: when the same item arrives from several feeds with different part numbers and no shared GTIN, the model sees several uncertain items and cites none of them. Whether a thin file like [llms.txt](https://geotoolbox.ai/blog/llms-txt) helps is a smaller question; the reachable page and the accurate feed are what actually matter. ## Can an Agent Understand Your Products? Being reachable gets you in the room. Being understood gets you considered. Agents do not read your page the way a person does; they extract structured data and act on it, so the gap between what your page shows a person and what it states in machine-readable form is where you lose. The consensus across every serious player is the same: move from human-readable to machine-readable. That means [schema markup](https://geotoolbox.ai/blog/schema-markup-for-ai) in JSON-LD, GS1 standards, and full attributes (price, brand, variants, sizing, availability, sustainability) in a form an agent can parse without guessing. Different vendors call it different things; it is the same instruction. Two practical points matter more than the acronyms. First, freshness: agents drop a recommendation when price or stock turns out wrong, so batch-hourly feeds are a liability. Second, the hard one: agents reward accuracy and consistency, not visual flair. If your edge is craftsmanship or sustainability rather than price, encode it into structured signals the model can weigh, or it flattens you to a spec sheet and picks someone cheaper. This is the same discipline as [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), pointed at the purchase instead of the citation. ## Will an Agent Trust and Recommend You? (And How to Check) Discovery and structure get you eligible. Being recommended is the prize, and it is the part the enterprise playbooks skip. The question brands actually ask, in their own words: why does the agent recommend an inferior competitor while my better, cheaper, in-stock product gets skipped? Sometimes it is the structured-data gap from the last section. Often it is trust signals the model can verify: reviews, ratings, return and warranty terms, consistent specs, even structured sentiment that lets it warn a shopper a shirt "runs small." Get those right and the agent keeps recommending you; get them wrong and it quietly moves on. The deeper problem is that the engines do not show you this. No native dashboard inside ChatGPT or Gemini tells you how often agents recommend you versus a competitor, or why you lost the comparison, so that signal has to come from outside the engines. That blind spot is why agentic visibility feels like guesswork. You cannot fix a recommendation you cannot see. Be careful what you assume carries over. Your backlink profile and domain authority, the moat you spent years on, count for far less here. Models lean on third-party best-of lists, reviews, and awards, and a given answer names only a few brands, closer to winner-take-all than the ranked page of links you know. In our experience auditing brands for AI visibility, most have never run the check, and the first scan is usually a surprise: a competitor named in the exact spot they assumed they owned. The fix is to treat it like rank tracking for a new surface. [Measure your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) across engines, watch your [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) on the prompts that matter, and find the [places where AI cites a competitor](https://geotoolbox.ai/features/citation-interceptor) instead of you, so you know where the recommendation is leaking. ## Can an Agent Actually Buy From You? This is where the hype and the reality split, and it pays to be precise because most write-ups are months out of date. ### What's Actually Live The headline version, "you can buy things inside ChatGPT," is now only half the story (the picture below is mid-2026). OpenAI launched [Instant Checkout in ChatGPT](https://openai.com/index/buy-it-in-chatgpt/) in September 2025 on the Agentic Commerce Protocol, starting with Etsy and Shopify merchants, then [pulled the native in-chat checkout back](https://searchengineland.com/chatgpt-instant-checkout-plan-change-471033) around March 2026 after only a tiny fraction of merchants went live: shoppers researched in ChatGPT but finished the purchase on the store. ChatGPT now leans discovery-first, comparing products and handing the shopper to the retailer's own checkout. Google went the other way: its [Universal Commerce Protocol](https://developers.googleblog.com/under-the-hood-universal-commerce-protocol-ucp/), launched at NRF in January 2026 with Shopify, Target, Walmart, Best Buy and others, runs end-to-end buying inside Gemini and Google's AI Mode. Google pushed it further at Google I/O in May 2026 with [Universal Cart](https://blog.google/products-and-platforms/products/shopping/google-shopping-cart/), a cross-retailer cart that follows a shopper across Search, Gemini, YouTube and Gmail. It adds buy-now-pay-later through Affirm and Klarna, brings in brands like Nike, Sephora, Ulta and Wayfair, and is rolling out beyond the US into Canada, Australia and the UK. So the real picture is competing rails, not one, and merchants are not listed automatically on any of them. They apply. ### The Protocols Doing the Plumbing Underneath, a stack of protocols does the plumbing, from OpenAI and Stripe's ACP to [Google's Agent Payments Protocol](https://techcrunch.com/2025/09/16/google-launches-new-protocol-for-agent-driven-purchases/), with the card networks now operating their own agent rails: [Visa's Trusted Agent Protocol](https://usa.visa.com/about-visa/newsroom/press-releases.releaseId.21716.html), which signs the agent's identity, and [Mastercard's Agent Pay](https://www.mastercard.com/global/en/news-and-trends/press/2025/april/mastercard-unveils-agent-pay-pioneering-agentic-payments-technology-to-power-commerce-in-the-age-of-ai.html), which hands the agent a scoped token instead of a raw card number.
ProtocolWhat it doesWho backs it
ACP (Agentic Commerce Protocol)Lets an agent complete a purchase with a merchant; separates discovery from checkoutOpenAI + Stripe
AP2 (Agent Payments Protocol)Cryptographically signed "mandates" that authorize an agent to pay a set amountGoogle, 60+ partners (Visa, Mastercard, PayPal)
UCP (Universal Commerce Protocol)Google's end-to-end standard: discovery through checkout inside AI surfaces like GeminiGoogle + Shopify, Target, Walmart, Etsy (20+ partners)
Trusted Agent ProtocolVerifies an agent's identity to the merchant with a cryptographic signature, so a known, trusted agent is recognized in the checkout flowVisa (Oct 2025)
Agent Pay"Agentic tokens" bind a tokenized card to a specific agent, merchant, and consent policy, so the model never holds the raw card numberMastercard (Apr 2025)
x402Instant agent payments in stablecoins over HTTPCoinbase
MCP (Model Context Protocol)Gives agents structured access to your data, like inventory and pricingAnthropic
Of these, Coinbase's x402 is drawing the most developer attention right now as a stablecoin settlement rail; we break down [what x402 is and how it works](https://geotoolbox.ai/blog/what-is-x402) separately. A second route skips the protocols entirely: agentic browsers like Comet and ChatGPT Atlas drive your existing site the way a person would, which makes "does my checkout work for a machine" its own question, covered in [is your site agent-ready](https://geotoolbox.ai/blog/agent-ready-website). And because Gemini's and Perplexity's checkout flows are still moving fast, treat any single vendor's claim as a snapshot. The smart play is to be ready on the open rails through your existing platform and keep [earning citations in ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt), since discovery now does most of the work. ## Your Agentic Commerce Readiness Checklist
![Six-gate agentic commerce readiness checklist, from reachable to measured.](/blog/agentic-commerce/agentic-commerce-readiness-checklist.png)
Six gates between your brand and an agent's recommendation — clear them in order.
You do not need to rebuild your stack. Clear the six gates in order, and most of the work overlaps with the GEO and technical SEO you should already be doing: 1. **Reachable.** AI crawlers can fetch your pages, and nothing critical hides behind JavaScript. 2. **Understandable.** Product schema in JSON-LD, GS1 standards, valid GTINs, full attributes. 3. **Accurate.** Price and stock are correct at query time, not on an hourly batch. 4. **Trustworthy.** Reviews, ratings, and return and warranty terms exposed and consistent. 5. **Payable.** Your platform supports agent-initiated checkout on the open rails like ACP and AP2. 6. **Measured.** You can see whether agents recommend you, and where they pick a competitor. If you do one thing this week, do the first and the last: confirm you are reachable, then check where you stand. ## What Agentic Commerce Means for Your SEO and GEO Work The anxieties underneath this are real. Brands worry about disintermediation, becoming a faceless supplier while the AI keeps the customer and the data. They worry about pay-to-play, the "discovery premium" McKinsey hints at, where being the agent's default becomes a paid placement the way ad slots did. And they worry the whole thing is hype. On disintermediation, the defensible move is to be both the recommended product and the merchant of record. You keep fulfillment, returns, the post-purchase relationship, and first-party data when the sale runs through your own agent-ready storefront, not only inside a platform's catalog. And not every channel plays the open game. Some marketplaces, including eBay and Amazon, restrict or block unauthorized third-party shopping agents to guard their own customer relationship, even as they push hard on agents of their own. Amazon [folded its Rufus assistant into Alexa for Shopping](https://www.cnbc.com/2026/05/13/amazon-ditches-rufus-ai-chatbot-in-favor-of-alexa-shopping-agent.html) in May 2026 and runs a "Buy for Me" feature that shops third-party sites on a customer's behalf, which drew retailer pushback because those retailers never authorized Amazon to complete purchases on their sites. Your site and the open rails are where you have control; marketplaces are a moving, sometimes hostile, target. The grounded read on the rest is calmer. Agentic commerce is not a new channel that replaces your SEO; it is the same core job (be found, be understood, be trusted) extended to the purchase. The structured data, clean feed, reviews, and citations you build to win [AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) are most of what gets you recommended too. The discipline carries over even where the old ranking signals do not, and for the first time you can measure it. If you have not mapped how the pieces fit, [GEO vs AEO vs SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) is the place to start. ## You Don't Have to Guess The brands that handle agentic commerce well are not the ones with the biggest budgets. They are the ones who looked. Most never check whether agents can reach their pages, read their products, or recommend them over a competitor, so they find out the slow way, when the traffic moves and the cause is invisible in their analytics. You can close that gap in an afternoon. geotoolbox is built for exactly this: a free [AI-readiness score](https://geotoolbox.ai/tools/ai-readiness) tells you whether agents can reach and read your site, and the visibility tools show where AI recommends you, and where it does not. ## Frequently Asked Questions ### What is the difference between agentic commerce and agentic AI? [Agentic AI](https://geotoolbox.ai/blog/agentic-ai) is the general capability: software that can plan and take actions on its own. Agentic commerce is that capability applied to shopping, payments, and merchants, where an agent finds products, compares them, and completes a purchase for a person. ### How big is the agentic commerce market? Estimates vary widely because the term is defined inconsistently. [Grand View Research](https://www.grandviewresearch.com/industry-analysis/agentic-commerce-market-report) put the market at $5.71 billion in 2025, growing to $65.47 billion by 2033, while McKinsey projects agents could orchestrate up to $1 trillion in US B2C retail and $3 to $5 trillion in global goods spend by 2030. Treat the long-range figures as direction, not precision. ### Is ChatGPT Instant Checkout still live? Not in its original form. OpenAI launched native in-chat Instant Checkout in September 2025 and pulled it back around March 2026 after it struggled to scale. ChatGPT now shows and compares products and sends shoppers to the retailer's own checkout via the Agentic Commerce Protocol (Etsy and Shopify merchants), while Google's rival Universal Commerce Protocol runs similar buying inside Gemini. ### Will I have to pay to be recommended by AI agents? For now, ChatGPT shows product results as organic and unsponsored, and merchants pay a small fee only on completed purchases, not for placement. Whether a paid "discovery premium" emerges later is an open question, and a reason to build organic agent visibility now, while it is still earned rather than bought. ### How do I check whether AI agents recommend my brand? Run the prompts a customer would and see who gets named, then track it across engines over time. A tool that monitors your AI share of voice and flags where AI cites a competitor instead of you turns a one-off spot check into something you can manage. ### What is an example of agentic commerce? You tell an assistant "reorder my coffee and find a grinder under $80 with good reviews." It checks your history, compares grinders across retailers, picks one, and either buys it or hands you a one-tap checkout. You may never open a product page. ## Sources - Buy it in ChatGPT: Instant Checkout and the Agentic Commerce Protocol - OpenAI - `openai.com/index/buy-it-in-chatgpt` - OpenAI's ChatGPT Instant Checkout plan just changed - Search Engine Land - `searchengineland.com/chatgpt-instant-checkout-plan-change-471033` - Google launches new protocol for agent-driven purchases (AP2) - TechCrunch - `techcrunch.com/2025/09/16/google-launches-new-protocol-for-agent-driven-purchases` - Agentic Commerce Market Size & Share Report - Grand View Research - `grandviewresearch.com/industry-analysis/agentic-commerce-market-report` - The agentic commerce opportunity - McKinsey - `mckinsey.com/capabilities/quantumblack/our-insights/the-agentic-commerce-opportunity-how-ai-agents-are-ushering-in-a-new-era-for-consumers-and-merchants` - What Is Agentic Commerce? - IBM - `ibm.com/think/topics/agentic-commerce` - Under the Hood: Universal Commerce Protocol (UCP) - Google Developers - `developers.googleblog.com/under-the-hood-universal-commerce-protocol-ucp` - 43% of Retailers Are Piloting AI Shopping Agents - PYMNTS - `pymnts.com/artificial-intelligence-2/2026/43-percent-of-retailers-are-piloting-ai-shopping-agents` - Google Shopping introduces Universal Cart and agentic shopping (Google I/O, May 2026) - Google - `blog.google/products-and-platforms/products/shopping/google-shopping-cart` - Visa introduces the Trusted Agent Protocol - Visa, October 2025 - `usa.visa.com/about-visa/newsroom/press-releases.releaseId.21716.html` - Mastercard unveils Agent Pay - Mastercard, April 2025 - `mastercard.com/global/en/news-and-trends/press/2025/april/mastercard-unveils-agent-pay-pioneering-agentic-payments-technology-to-power-commerce-in-the-age-of-ai.html` - Amazon ditches Rufus chatbot for Alexa shopping agent - CNBC, May 13, 2026 - `cnbc.com/2026/05/13/amazon-ditches-rufus-ai-chatbot-in-favor-of-alexa-shopping-agent` - Agentic Commerce Market Impact Outlook - Morgan Stanley - `morganstanley.com/insights/articles/agentic-commerce-market-impact-outlook` --- ## E-E-A-T for AI Search: What Actually Makes Content Citable > E-E-A-T is not a ranking factor or a score. But the trust signals behind it decide whether ChatGPT, Perplexity, and AI Overviews cite you. Here is what works. - Canonical: https://geotoolbox.ai/blog/eeat-ai-search - Published: 2026-06-24 · Updated: 2026-08-16 You added the author bio. You earned the rankings. And ChatGPT, Perplexity, and Google's AI Overviews still never mention you. If E-E-A-T is supposed to be your ticket into AI search, why does doing it by the book change nothing? E-E-A-T for AI search is widely misunderstood. It is not a score you raise or a ranking factor you switch on. It is a label for the trust signals AI engines reach for when they decide whose content to cite. This is about which signals move that decision, which ones are theater, and what to do when you are doing everything right and still going uncited. ## What E-E-A-T Actually Is E-E-A-T stands for Experience, Expertise, Authoritativeness, and Trustworthiness. It is the language Google's [Search Quality Rater Guidelines](https://geotoolbox.ai/glossary/e-e-a-t) use to describe credible content. Google added the second E, for Experience, in [December 2022](https://developers.google.com/search/blog/2022/12/google-raters-guidelines-e-e-a-t), to value first-hand knowledge alongside formal expertise. One detail matters more than the acronym: the four are not equal. Google states plainly that "trust is most important," and that the other three exist to support it. A page that is not trustworthy has low E-E-A-T no matter how experienced or expert it looks. The framework gets extra weight on [Your Money or Your Life (YMYL)](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) topics, where bad information can cause real harm. None of these are things an engine reads directly, and that is the part most guides skip. They are human-rater concepts, and both Google's ranking systems and AI answer engines have to infer them from signals they can observe. That gap, between the idea and the observable signal, is where most E-E-A-T advice goes wrong.
PillarWhat Google means by itWhat an AI engine can observe
ExperienceFirst-hand, real-world use of the topicOriginal photos, data, dated results, and first-person detail it can extract
ExpertiseKnowledge, skill, or credentialsA named author with a machine-readable identity and matching profiles elsewhere
AuthoritativenessReputation as a go-to sourceOther credible sites citing, linking to, and mentioning you
TrustworthinessAccuracy, transparency, and safetyConsistent facts across the web, clear sourcing, HTTPS, real contact details
Read the right-hand column again. Most of those signals, especially authority and trust, are things other people say about you, not things you can declare on your own page. Hold onto that, because it explains most of how AI citation works. ## Is E-E-A-T a Ranking Factor? No, and That Matters More for AI There is no E-E-A-T score, and it is not a ranking factor. This is [Google's own position](https://developers.google.com/search/docs/fundamentals/creating-helpful-content): "While E-E-A-T itself isn't a specific ranking factor, using a mix of factors that can identify content with good E-E-A-T is useful." Quality raters score sample results to check whether the ranking systems are working. Their ratings do not flow back as a number attached to your page. So when a tool offers to grade your "E-E-A-T score," it is inventing a metric Google says does not exist. The same goes for the checklist version of the idea, the one where you bolt on an author box and a "reviewed by" line and expect rankings to move. That is a costly myth. Google's Danny Sullivan has been blunt that [author bios are not a ranking signal](https://searchengineland.com/google-eeat-misconceptions-437445): having an expert write something does not magically make it rank, and adding a credentials line does not flip a switch. The bio still matters, but not as a lever you pull on your own page. It matters as evidence an engine can connect to a real, recognized person. AI answers raise the stakes. The engines are even further from your page than Google's crawler is: they cannot see your intentions or your effort, only signals. So the same correlated signals that nudge search rankings weigh even heavier when an engine decides what to cite. A page it cannot place or trust rarely surfaces in an answer at all. ## Does E-E-A-T Apply to ChatGPT, Perplexity, and AI Overviews? Strictly, no. E-E-A-T is Google's vocabulary for its human raters. ChatGPT, Claude, and Perplexity do not run a function called E-E-A-T, and anyone who tells you they do is guessing. The better question is whether the same underlying signals decide who gets cited. There, the answer is yes. Many AI answers are built with retrieval: the engine runs a search, pulls a set of pages, and writes from them. [How AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) under the hood is a retrieval step feeding a generation step. That retrieval step leans on the same web that search engines have already sorted by quality, so the trust signals you would build for E-E-A-T are largely the ones that get you into the retrieved set. Google's AI Overviews lean on pages Google's core systems already rank, though less tightly than they used to. Ahrefs put the share of AI Overview citations coming from the top 10 at [76% in July 2025](https://ahrefs.com/blog/eeat-seo/), then at [38% in a larger March 2026 study](https://ahrefs.com/blog/ai-overview-citations-top-10/). Ranking still gets you considered; it no longer guarantees the citation. They are not, however, the same thing. A strong rank gets you considered, not chosen. Plenty of pages rank well and never get named in an answer, and plenty of cited sources sit far outside page one. So treat ranking as the price of entry and citation as a second selection on top of it, one that weighs trust and corroboration even harder. The next section is about why that second selection so often goes against the page that "did everything right." ## Why You Rank #1 and Still Do Not Get Cited The single most common frustration we hear: "I rank first on Google, my page has an author bio and schema, and ChatGPT still never names me." The reason is that AI engines lean heavily on trust at the level of the entity, not just the page. A page can be excellent in isolation. But an engine deciding whether to cite a claim wants corroboration, and corroboration lives off your site. It is looking at whether other credible sources talk about you, link to you, and agree with you. The shorthand: AI does not pull its read of you from what your homepage says, it pulls from what the rest of the web says about you. That is why a thin Reddit thread can beat your detailed guide. Reddit is a recognized entity that thousands of independent voices reinforce, and one the engines have struck content deals with, so it is an easy thing to quote. Your better page, from a domain the engine has barely seen mentioned anywhere, is a riskier bet. It is not fair, but it is legible. The data points the same way. Ahrefs found that the off-site signals correlating most strongly with AI Overview mentions were [branded web mentions, at 0.664 correlation](https://ahrefs.com/blog/eeat-seo/), ahead of branded anchors and branded search volume. And a large share of what AI cites is structurally out of your hands: Ahrefs also found that [roughly two-thirds of ChatGPT's most-cited sources are off-limits to marketers](https://ahrefs.com/blog/chatgpts-most-cited-pages/), dominated by Wikipedia, reference sites, and other pages you cannot easily influence through outreach. So the work that moves AI citation is not another on-page tweak. It is becoming an [entity the engines recognize](https://geotoolbox.ai/blog/entity-seo): consistently named, consistently described, and corroborated across the [sources an AI engine already trusts](https://geotoolbox.ai/glossary/ai-citation). Your page is where you convert that trust into an answer, not where you create it. ## The Signals AI Reads, and the Theater It Ignores Once you accept that engines infer trust from observable signals, the to-do list sorts itself into two piles. One pile changes whether you get cited. The other is busywork that feels like E-E-A-T because it uses the vocabulary.
![Five trust signals AI engines can observe, and three E-E-A-T theater items they ignore.](/blog/eeat-ai-search/ai-trust-signals-vs-theater.png)
The two piles: signals engines can observe and corroborate, versus busywork that only borrows the E-E-A-T vocabulary.
Moves the needleTheater (feels like E-E-A-T, mostly is not)
A named author with a real, linkable identity (Person schema plus a sameAs to profiles that exist)A "reviewed by Dr. X" badge with no verifiable person behind it
Off-site brand mentions and earned media on sources engines already trustSchema markup treated as a trust score on its own
Answer-first passages an engine can lift without contextAn llms.txt file expected to boost citations on its own
Original first-hand data, tests, and dated screenshotsPrecise vendor stats like "answer-first earns 67% more citations"
Pages AI crawlers can fetchStacking author boxes on pages no one off-site corroborates
Two items on the theater side surprise people. Schema is one. It makes your author and organization machine-readable, which genuinely helps an engine connect the dots, but it is a signal-exposer, not a trust score. When Ahrefs tracked [1,885 already-cited pages that added schema](https://ahrefs.com/blog/schema-ai-citations/), it found no meaningful citation lift. Use schema to expose facts that are already true, not to manufacture authority. The other is the genre of precise citation statistics. The "answer-first gets cited 67% more" numbers trace back to vendor blog posts with no published method. The durable findings are blunter: be trustworthy, be original, be extractable, be corroborated. When the [craft of writing pages LLMs cite](https://geotoolbox.ai/blog/ai-content-optimization) is done well, the citations follow without a magic percentage attached. ## Experience vs Expertise: the "E" Most Content Fakes The two front E's are not synonyms, and the difference is the part most content gets wrong. Expertise is theoretical knowledge: a doctor who can list the symptoms of the flu. Experience is lived: a patient describing what the flu felt like. Google added Experience because for a lot of queries, readers want the second kind, and a credentialed summary of what everyone already knows is not it. This is also the hardest signal to fake, which is exactly why it is worth investing in. You cannot bolt on first-hand experience with a byline. It shows up as the texture only a real user produces: the specific number, the screenshot with your own data in it, the thing that went wrong that no summary would mention, the dated before-and-after. Engines and readers both reward that detail because it cannot be cheaply synthesized. In our experience auditing pages that get cited versus pages that do not, the citable ones almost always contain something the writer could only know by doing the thing. The uncited ones read like a competent rewrite of the top ten results, which is precisely what a model can already generate, so it has no reason to send anyone there. Can engines tell the difference between real experience and credential cosplay? Not perfectly, but they lean on corroboration to approximate it, and content that mimics expertise without any external validation tends to be treated as exactly that. The reliable move is not to perform experience. It is to have some, then make it extractable. ## Reachability: the Gate Before Any of This Matters There is one signal that sits in front of all the others, and it is the one teams forget. An engine that retrieves the web cannot cite a page it cannot fetch. If your robots.txt blocks the [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) that feed those answers, the trust work never gets read, because the page never enters the retrieval set. It is a common silent failure in AI visibility, and one of the easiest to fix. OpenAI, Anthropic, and Perplexity each run named crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) you can allow independently, and it is easy to find one quietly disallowed by a default rule or a security plugin nobody revisited. [Google's AI Overviews](https://geotoolbox.ai/blog/what-are-google-ai-overviews) are the exception: they ride on ordinary Googlebot, so there it is your normal crawl and index settings that decide inclusion, not a separate AI crawler. The free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows which AI crawlers your robots.txt allows or blocks, with the exact line to fix. Run it before you invest in any of the trust work above. Reachability does not earn you a citation. It just makes you eligible for one. Treat it as the gate it is: confirm it first, then go build the trust that gets you through. ## How to Measure E-E-A-T for AI Without a Paid Tool Since there is no E-E-A-T score, you cannot check a number. But you can check the thing the score is a proxy for: what the web, and the engines reading it, currently believe about you. Most advice here points straight at a paid toolkit. You do not need one to get started. The free version takes ten minutes. Search your brand and your key topics while excluding your own domain (a `your brand -site:yourdomain.com` query), so you only see what other sources say. Then ask the engines directly: pose the questions your customers ask to ChatGPT, Perplexity, and Google's AI Overviews, and watch who gets named. You are looking for three things: whether you appear at all, who describes you when you do, and whether the description is accurate. Silence usually means you have not cleared the entity bar yet, though it can also be query variance or a topic that does not trigger an AI answer. A wrong description means the corroboration exists but points the wrong way. Google hands you a self-check for this without naming it E-E-A-T. Its [helpful-content guidance](https://developers.google.com/search/docs/fundamentals/creating-helpful-content) asks Who created the content (is there a real, identifiable author), How it was made (is any automation disclosed), and Why it exists (to help a reader, or just to rank). Those three questions map onto what an AI engine can verify: a named author it can resolve, content whose provenance is clear, and a purpose that is not obviously manipulation. If your pages cannot answer Who, How, and Why, an engine cannot either. That manual pass tells you where you stand. The structured version is what our [AI Readiness](https://geotoolbox.ai/tools/ai-readiness) check automates, and tracking it over time is how you turn a one-off look into [a way to measure AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) as you build. Either way, the point is the same: stop trying to grade a page, and start watching whether engines can find, trust, and correctly describe you. ## How the Engines Differ No single playbook performs identically everywhere, because the engines retrieve and weigh sources differently. None of them publish a trust formula, so read the table below as observed tendencies, not settled factors.
EngineWhere it pulls answers fromWhat it appears to over-index on
Google AI OverviewsPages Google already ranks (ordinary Googlebot)The same authority signals that drive Google rankings
ChatGPT (search)Its own search index plus the model's training memoryEstablished, widely-referenced sources and familiar brands
PerplexityIts own crawl plus live fetchesRecent pages and several independent sources agreeing
GeminiGoogle's index and groundingRecognized entities Google already connects in its Knowledge Graph
One caveat the table hides: an engine can name a well-known brand straight from training memory, with no retrieval at all. That is the long game of being an entity. You become part of what the model already knows, not just what it can look up. ## A Realistic On-Ramp for New or Small Sites If your site is new and has no link history, "build authority" can sound like "already be famous." You cannot conjure years of brand mentions overnight, but you can pick the moves that compound fastest, including a focused plan for earning links from relevant sites through the right [link building platforms](https://geotoolbox.ai/blog/best-link-building-platforms). Lead with depth, not breadth. One subject covered thoroughly, as a connected cluster rather than a single page, builds [topical authority](https://geotoolbox.ai/blog/topical-authority) an engine can recognize faster than scattered posts across ten topics. Pair that with the one thing a bigger competitor cannot copy: your own first-hand data, your own tests, your own numbers. Originality is the cheapest entity signal a small site can create, because it is real. Then get named where your audience already talks, in communities, on podcasts, in roundups, anywhere an independent source can mention you. A handful of credible mentions does more for AI citation than another self-published page. If you work in a Your Money or Your Life field, health, finance, or legal, the bar is higher. Engines lean harder on named, credentialed authors and on outside corroboration, and they filter fringe claims more aggressively, so first-hand proof and recognized sourcing matter even more. The [full sequence for optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) walks through the order. The short version: reachable first, original and extractable second, corroborated third. ## Frequently Asked Questions ### Is E-E-A-T a ranking factor? No. Google states that E-E-A-T is not a specific ranking factor and that there is no E-E-A-T score. It is a framework its human quality raters use to judge content, and the rater feedback helps Google tune its ranking systems over time. You optimize for the signals that correlate with it, not for a number. ### Does E-E-A-T matter for ChatGPT and Perplexity? Those engines do not run Google's E-E-A-T framework, so not literally. But they retrieve sources from a web already filtered by search and reach for the same underlying trust signals, so the work you do for E-E-A-T is also what makes you eligible to be cited in AI answers. ### Why does AI cite Reddit over my detailed expert content? Because AI engines weigh trust at the entity level and prefer corroboration. Reddit is a recognized source that thousands of independent voices reinforce, which makes it an easy thing to quote. A stronger single page from a domain the engine rarely sees mentioned elsewhere is one it has less reason to trust, even when it is better. ### Does schema markup improve AI citations? Schema makes your author and organization machine-readable, which helps an engine connect your content to a known entity. But it is not a trust score on its own. Ahrefs tracked [1,885 already-cited pages that added schema](https://ahrefs.com/blog/schema-ai-citations/) and found no meaningful citation lift, so use schema to expose facts that are already true, not to manufacture authority. ### Does using AI to write hurt E-E-A-T? Not by itself. Google says it judges content on quality, not on whether a human or AI produced it, and only treats AI made mainly to manipulate rankings as spam. The real risk is subtler: AI is very good at the generic rewrite of the top results, and that is exactly the content an engine has least reason to cite. Use it to draft and edit, then add the first-hand experience and original data it cannot invent. ### Can you fake E-E-A-T, and will AI catch it? You can add the surface markers (an author box, a "reviewed by" line, a credentials list), but without off-site corroboration they do little. Engines approximate genuine experience and authority through external validation, and content that mimics expertise with nothing backing it up tends to be treated as exactly that. ## Stop Grading the Page E-E-A-T for AI search is not a checklist you complete or a score you raise. It is whether the engines can reach your content, trust the entity behind it, and find your claim corroborated somewhere they already believe. Get those three right and you stack the odds for citation. Stack author boxes on an unreachable, uncorroborated page and you will not, no matter how the page scores in some tool. If you are not sure which of the three is holding you back, start by measuring instead of guessing. Our free [AI Readiness](https://geotoolbox.ai/tools/ai-readiness) check shows where you stand on reachability and the trust signals AI engines actually read, which is the difference between optimizing a page and earning a citation. geotoolbox is the AI bot debugger we built for exactly this: seeing your site the way an answer engine does, before you spend a month on signals it never gets to use. ## Sources - Google Search Central, Creating Helpful, Reliable, People-First Content - `developers.google.com/search/docs/fundamentals/creating-helpful-content` - Google Search Central Blog, Our latest update to the quality rater guidelines: E-A-T gets an extra E for Experience - `developers.google.com/search/blog/2022/12/google-raters-guidelines-e-e-a-t` - Search Engine Land, Debunking common Google E-E-A-T misconceptions - `searchengineland.com/google-eeat-misconceptions-437445` - Ahrefs, E-E-A-T: How to Build Trust and Boost Web & AI Visibility - `ahrefs.com/blog/eeat-seo` - Ahrefs, What Are ChatGPT's Most Cited Pages? - `ahrefs.com/blog/chatgpts-most-cited-pages` - Ahrefs, We Tracked 1,885 Pages That Added Schema Markup - `ahrefs.com/blog/schema-ai-citations` --- ## Grok 5: Release Date, Specs & What's Confirmed (August 2026) > Grok 5 isn't out yet. A dated look at xAI's next model: the slipping release date, confirmed-vs-rumored specs, the AGI claim, and what it means for AI search. - Canonical: https://geotoolbox.ai/blog/grok-5 - Published: 2026-06-24 · Updated: 2026-08-13 Grok 5 is the most-anticipated AI model that has not actually shipped. As of August 2026 it is still in training, xAI has not committed to a real release date, and almost everything written about its specs is rumor dressed up as fact. Here is what is actually confirmed, what is guessed, and what it means for you. Most "Grok 5" coverage either breathlessly announces features that do not exist yet or recycles a release date that already slipped. What follows separates the confirmed facts from the hype, flags the numbers nobody can verify, and skips the assumption that a bigger model is automatically a better one. One quick disambiguation first: this is about Grok 5, xAI's next flagship model, not "level 5" in Grok's Ani companion, which is a different thing entirely. ## Is Grok 5 Out Yet? The Release Date and the Slipping Timeline No. As of August 2026, **Grok 5 has not been released.** xAI's next flagship model is still in training, and the only thing the company has officially said about it is a single line in its January 2026 funding announcement confirming the model exists and is being trained. There is no model card, no spec sheet, and no public benchmark. Beyond that in-training status and a handful of corroborated facts, almost everything specific is reporting, a leak, or Elon Musk posting on X. The release date has slipped more than once. Musk first pointed at late 2025. That moved to Q1 2026, a target he gave in November 2025, then to a Q2 2026 window xAI referenced after Q1 came and went. Both windows have now passed with no launch: Q2 closed June 30 and, as of early July, Grok 5 is still training. The prediction market Polymarket, where traders bet real money on the release date, tells the same story: its "Grok 5 released by June 30, 2026" contract was priced as a real possibility earlier in the year and expired a long shot as that deadline came and went. What xAI shipped instead: on July 8, 2026 it released [**Grok 4.5**](https://geotoolbox.ai/blog/grok-4-5), an "Opus-class" flagship for coding and agentic work that Musk described as "roughly comparable to Opus 4.7, but much faster." Reportedly built on xAI's V9 foundation and trained partly on Cursor session data, it offers a 500K-token context window at $2 per million input tokens and $6 per million output tokens, and landed 4th on the Artificial Analysis Intelligence Index (EU-blocked at launch under the AI Act, it became fully available across the EU on July 16, 2026). [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) has since taken over as the interim flagship (August 12, 2026); neither it nor Grok 4.5 is Grok 5. Why the delays? Partly because training a frontier-scale model is genuinely slow and hard. But also because xAI spent the first half of 2026 shipping nearly everything except Grok 5: Grok Voice, the Grok Imagine video model, point releases up to **Grok 4.5** and now **Grok 4.6** (the current flagship, per [xAI's own model docs](https://docs.x.ai/docs/models)), and a separate 1.5-trillion-parameter coding model (Grok V9-Medium, covered below). Musk also has a long record of optimistic AI and Tesla timelines that land late, which is why experienced observers treat every Grok 5 date as a guess until xAI posts otherwise. Underneath the slip sit two harder problems, unsolved engineering bets rather than mere optimism: networking a multi-trillion-parameter mixture-of-experts across a gigawatt-scale cluster strains the interconnect, and training on X's real-time firehose risks feeding the model its own AI-generated output. Compute itself is not the bottleneck; xAI has poured resources into its Colossus 2 supercluster since [SpaceX acquired the company in February 2026](https://techcrunch.com/2026/02/02/elon-musk-spacex-acquires-xai-data-centers-space-merger/). The practical move is to watch the source rather than the hype. xAI announces real launches on its own News page, the official @xAI account on X, and the model docs that list what is live right now. Our own [guide to Grok and its model history](https://geotoolbox.ai/blog/what-is-grok) covers up to [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) for the same reason: that is the newest model you can actually use today. Concrete signs Grok 5 is close, rather than another rumor cycle: - a `grok-5` entry appearing in the [xAI model docs](https://docs.x.ai/docs/models) - an official model card or system card - a benchmark or leaderboard entry, such as LMArena or SWE-Bench - a launch post from the @xAI account, not a one-line Musk reply ## Grok 5 Specs: What's Confirmed vs What's Rumored The split is straightforward. The table below separates what xAI has actually confirmed from the numbers that circulate as if they were official. The pattern worth noticing is that almost every specific spec sits in the rumored column.
![Status table separating the one confirmed Grok 5 fact from rumored and unconfirmed claims.](/blog/grok-5/grok-5-confirmed-vs-rumored.png)
The Grok 5 claims in circulation, sorted by what xAI has actually confirmed — which is almost nothing.
Claim about Grok 5StatusWhat we actually know
In training on Colossus 2ConfirmedStated in xAI's January 2026 funding announcement; Colossus 2 is xAI's Memphis supercluster
Parameter count of 6 to 10 trillion (mixture-of-experts)RumoredMusk and leaks float figures from roughly 6T to 10T; reporting disagrees on which is Grok 5 versus a sibling model training alongside it on Colossus 2. No model card confirms any number
A 1.5 million token context windowEstimateNot official; the shipping Grok 4.3 is 1M tokens and the Grok Build coding model is 256K
Native multimodal and live X dataExpectedConsistent with Grok's existing direction, but not specified for Grok 5
A "Reality Engine" fact-checkerUnconfirmedA label coined in blog posts, not a feature xAI has announced
A confirmed release dateNoneThe late-2025, Q1 2026, and Q2 2026 targets all passed with no launch (Q2 closed June 30)
Two things drive the confusion. There is no Grok 5 model card, so reporters fill the gap with leaks and Musk's off-the-cuff comments, and those harden into "facts" through repetition. And xAI is training several models at once on Colossus 2 at different sizes, so a parameter figure that belongs to a different run gets pinned on Grok 5. The parameter number deserves the most caution, because it is the stat most likely to mislead. Grok 5 is described as a [mixture-of-experts model](https://geotoolbox.ai/glossary/mixture-of-experts), which means that even if the total really runs into the trillions, only a fraction of those parameters fire on any given query. A multi-trillion-parameter mixture-of-experts model is not automatically smarter than a smaller dense one; it is a bigger library where each question still pulls only a few books off the shelf. More capacity helps, but it does not convert linearly into capability, and it says nothing about reasoning quality, which is where current models actually compete. The context window is the other figure to handle carefully. The widely repeated 1.5 million tokens is an estimate, not a confirmed spec. For reference, xAI's shipping models list a [1 million token context window](https://geotoolbox.ai/glossary/context-window) for Grok 4.3 and 256K for the Grok Build coding model. Grok 5 may extend that, but no official number exists yet. Treat "6 trillion parameters" the way you would treat an engine's displacement. It tells you something about scale, not about how the thing drives. The numbers that will actually matter (reasoning benchmarks, coding accuracy, hallucination rate) do not exist for Grok 5 yet, because the model has not shipped. ## Grok 5 vs Grok V9-Medium: Why People Confuse Them If you have seen a headline like "Elon unveils Grok 5 with 1.5 trillion parameters," ignore it. That is not Grok 5. It describes **Grok V9-Medium**, a separate, coding-focused model that completed its base training in mid-2026 (with public release expected around then), reportedly trained partly on data from the Cursor code editor to be strong at programming. It is roughly 1.5 trillion parameters. Grok 5, the model this article is about, is the much larger flagship rumored at around 6 trillion, and it is still in training. The two get merged for a simple reason: xAI runs two parallel naming systems and rarely explains which is which. There are the consumer release names you see in the app, Grok 4, Grok 4.1, Grok 4.3, and there are internal foundation-model version numbers like "V9." They do not line up cleanly, so a "V9" coding model and "Grok 5" sound like the same generation when they are not. If you are deciding whether to wait for Grok 5, you do not want to base that decision on a separate coding model that is nearly out. And when you read a benchmark claiming "Grok 5 tops the coding charts," check whether the number actually came from V9-Medium, the model purpose-built for code. One more name to clear up while we are here: "level 5" on Grok has nothing to do with the model. That phrase refers to an affection tier in Grok's "Ani" companion persona, a separate feature entirely. Searching "Grok 5" and landing on "Grok Ani level 5" content is a common wrong turn, and the two have nothing in common beyond a number. ## What Grok 5 Is Expected to Do Strip out the hype and a consistent picture of xAI's direction remains, even with the Grok 5 specifics unconfirmed. Versus the Grok 4.3 you can use today, the expected jumps fall into four areas. **Bigger multi-agent reasoning.** Grok's 4.20 beta (released February 2026, an earlier build than 4.3 despite the bigger-looking number) introduced a setup where several specialized [AI agents](https://geotoolbox.ai/glossary/ai-agent) work a problem in parallel and cross-check each other. Grok 5 is widely expected to scale that up. The caveat: more agents is an architecture choice that adds cost per answer, and it does not guarantee a quality jump. **Stronger real-time grounding.** Grok's defining trait is live access to public posts on X and the open web, which lets it answer about something that happened minutes ago. A bigger model on the same live pipe is the most predictable upgrade. **Deeper multimodality.** xAI already ships image and video generation ([Grok Imagine](https://geotoolbox.ai/blog/grok-imagine)) and voice as separate features; folding native video and audio understanding into one model is the expected next step. **More agent, less chatbot.** Musk frames Grok 5 as moving beyond a simple chatbot toward a system that plans and executes multi-step tasks. Every major lab is pushing that direction, so execution will matter more than the pitch. Two widely repeated "features" deserve a flag. The "Reality Engine," an automatic fact-checker that verifies claims against X in real time, is a label coined in blog posts; xAI has never announced it. And "rapid learning," the idea that Grok 5 will improve weekly from user feedback, describes a deployment pattern xAI has only talked about loosely. Treat both as unconfirmed until xAI says otherwise. ## The AGI Claim: What Musk Said vs What's Likely The loudest claim about Grok 5 is that it might be artificial general intelligence. In October 2025, Musk posted that his "estimate of the probability of Grok 5 achieving AGI is now at 10% and rising," and added that "Grok 5 will be AGI or something indistinguishable from AGI," [as reported by Teslarati](https://www.teslarati.com/elon-musk-grok-5-now-has-10-percent-chance-of-becoming-worlds-first-agi/). Both lines traveled far. Two things are worth holding onto when you read them. First, "10%" is a self-stated probability, not a measured one, and it is an odd number to celebrate: a 10% chance of AGI is a 90% chance of not-AGI by Musk's own math. There is no agreed definition of AGI, no benchmark that declares it, and no peer-reviewed Grok 5 result to point to, because the model is not out. The claim is a posture, not a finding. Second, the people who build these systems are openly skeptical. When Musk made the claim, OpenAI research scientist Gabriel Petersson [joked](https://futurism.com/artificial-intelligence/openai-researcher-mocks-elon-musks-agi) that there was a "10 percent chance Elon declares he reached AGI a fourth time," a jab at a pattern of AGI declarations that never quite land. That captures the broader expert mood: interest in the scale, doubt about the label. There is also a serious technical argument that a bigger Grok may not be the leap it sounds like. On public leaderboards, recent frontier models cluster within roughly five to ten points of each other, which suggests raw scale is hitting diminishing returns and that reasoning techniques, rather than parameter counts, are where the gains now come from. A model at this scale could land as a strong, expensive, competitive system that still [hallucinates](https://geotoolbox.ai/blog/ai-hallucinations) and still trips on the same hard reasoning problems as everyone else. That is the likeliest outcome on current trends, and it would be a perfectly good model that simply isn't science fiction. The realistic read: Grok 5 will probably be very capable. "AGI" here is a marketing label, not a measured result, and you should price it accordingly. ## Grok 5 vs the Current Frontier Models The straight answer to "how does Grok 5 compare to GPT-5.6, Claude, and Gemini" is that you cannot compare it to anything yet, because it has no published results. What you can compare is its expected shape against what already ships today.
DimensionGrok 5 (expected, unconfirmed)What ships today
StatusIn training, no release dateGrok 4.6, GPT-5.6, Claude Opus 5, and Gemini 3.1 Pro are all live now
ParametersRumored 6T to 10T (mixture-of-experts)Not officially disclosed for most flagships, so any comparison is an estimate
Context windowEstimated ~1.5M tokensGrok 4.3 and several rivals already offer 1M tokens
Real-time dataLive X and web (Grok's signature edge)Grok's live access is a genuine edge; most rivals rely on slower web search or none
Proven benchmarksNone yetRivals have published reasoning, coding, and safety results
PriceUnannouncedGrok 4.3's API is $1.25 / $2.50 per million tokens, among the cheaper frontier options
A few things are safe to say without a launch. Grok's genuine, repeatable edge is that live connection to X and the open web, which makes it strong for breaking news and real-time sentiment in a way models without live retrieval are not. It also tends to be cheaper at high volume on the API. What is not safe to say is that Grok 5 will "beat everything," the framing in most hype coverage. Bigger has not reliably meant better for a while now, and its rivals are not standing still. Grok also carries reputational baggage a better benchmark won't erase: past content-moderation incidents and wariness of the Musk brand mean some buyers will pass on it whatever the scores say. On its own turf of live data and real-time grounding it stays genuinely strong; the gap shows up most on raw reasoning leaderboards. For a real head-to-head on the models you can actually use today, we keep two current comparisons updated: [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt) and [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude). Both pit shipping models against each other, which is the only comparison that means anything right now. So should you wait for Grok 5? For almost any real task, no. A model you can use today beats a better one you cannot, and if Grok 5 turns out to be the leap Musk promises, switching to it later costs you nothing. Waiting does. ## Grok 5 Pricing and Access: Will It Be Free? xAI has not announced Grok 5 pricing, so anything specific is a guess. What you can do is read the pattern from how xAI prices Grok today and assume Grok 5 slots in at the top of it. There will very likely be a limited free tier. Grok already offers free, rate-limited access on X, at grok.com, and in the app, and xAI has kept a no-cost entry point through its releases. Expect the same here: a capped free tier on the older models, with Grok 5 itself reserved for paying users at launch, the way new flagships usually roll out. For full access, the existing paid ladder is the guide: SuperGrok at around $30 a month, a heavier SuperGrok tier near $300 a month for priority access and the most compute-intensive features, and X Premium+ bundling Grok with the platform. Treat those as the current shape, not Grok 5's confirmed prices, and check the live numbers on xAI's site, because consumer pricing shifts often. Our [Grok pricing guide](https://geotoolbox.ai/blog/grok-pricing) covers the current tiers in detail. Developers have a firmer anchor. xAI's [official model docs](https://docs.x.ai/docs/models) list current API pricing at $1.25 per million input tokens and $2.50 per million output tokens for Grok 4.3, with the Grok Build coding model lower. xAI has priced its API aggressively against OpenAI and Anthropic, so a Grok 5 endpoint will likely cost more than 4.3 per token while staying competitive. Early analyst estimates float something like $5 to $8 per million input tokens and $20 to $30 on output, roughly four to six times Grok 4.3 on input and eight to twelve times on output, but those are projections, not announced prices. When it lands, the API model name will probably be straightforward, though even that is a guess until xAI lists it. ## What Grok 5 Means for AI Search and Your Brand Here is the part the release-date trackers skip. Whether or not Grok 5 turns out to be "AGI," a bigger, more capable Grok with live access to X means more people will ask Grok questions where your brand could come up, and more of them will treat its answer as the answer. That makes Grok part of your visibility surface, the way Google search results have been for years. The mechanics decide what you do about it. Grok answers in two ways: from what it absorbed during training, its parametric memory, and from what it retrieves live off X and the open web the moment you ask. What Grok and the other engines cite tends to lean on a familiar handful of sources: public posts on X, Reddit threads, YouTube, Wikipedia, and a small set of authoritative web pages. Grok leans on X more heavily than the others do, so a consistent presence on X and in the communities it surfaces matters more for Grok specifically than it does for, say, ChatGPT. If your brand is absent or thinly represented across those, the model fills the gap with whatever it can find, which is exactly how confident, wrong answers about companies happen. None of this waits for Grok 5. The work that makes you visible in a future Grok is the same work that makes you visible in the Grok shipping today, and in ChatGPT, Gemini, and Claude. It comes down to two things: being a complete, consistent entity across the open web so the model's training memory has accurate material to draw on, and being a clean, citable source so the live retrieval layer can find and quote you. That is the core of [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), and it holds across model launches in a way that chasing any single release does not. In our experience, the brands that take each new model launch in stride are the ones already measuring where they stand. They are not refreshing release-date trackers; they are watching what the engines actually say about them and closing the gaps. The first step is just seeing it: [tracking your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) across engines so a new model becomes a data point instead of a fire drill. ## The Bottom Line on Grok 5 Grok 5 is real, it is in training, and it is late. The 6-trillion-parameter spec, the 1.5-million-token context, and the AGI headline are all unconfirmed, and the most likely outcome is a powerful, expensive, competitive model rather than a machine that rewrites the definition of intelligence. Watch xAI's own channels for the real date, and treat unconfirmed "Grok 5 is here" coverage as speculation until then. What should not wait is your own visibility. A bigger Grok only raises the stakes on whether AI engines describe your brand accurately. The [Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) shows you what ChatGPT, Gemini, Claude, Grok, and the other engines cite when your topic comes up, and where your brand is missing from the answer, so you are ready for Grok 5 whenever it finally ships. ## Frequently Asked Questions ### When is Grok 5 coming out? There is no confirmed release date as of August 2026. xAI originally targeted late 2025, then Q1 2026, then a Q2 2026 window; all three passed with no launch, and Q2 closed June 30. The model is still in training, and prediction markets steadily lowered the odds of a first-half-2026 launch. In the meantime, xAI shipped the interim Grok 4.5 flagship on July 8, 2026, then Grok 4.6 on August 12, not Grok 5. Watch xAI's News page and the @xAI account for the official word rather than third-party predictions. ### Is Grok V9-Medium the same as Grok 5? No. Grok V9-Medium is a separate, roughly 1.5-trillion-parameter coding model that finished its base training in mid-2026. Grok 5 is the larger flagship, rumored at around 6 trillion parameters, that is still in training. Headlines merge the two because xAI's internal version numbers do not match its consumer release names. ### What is "level 5" on Grok? That is unrelated to the Grok 5 model. "Level 5" refers to an affection tier in Grok's "Ani" companion persona, a separate feature. If you searched for the model and found companion content, that is the mix-up. ### How many parameters does Grok 5 have? No official figure exists, because xAI has not published a Grok 5 model card. Musk and leaks float figures from roughly 6 trillion to 10 trillion in a mixture-of-experts design, and reporting disagrees on which number is Grok 5 versus a sibling model training alongside it. Treat all of them as unconfirmed. ### Will Grok 5 achieve AGI? Almost certainly not in any rigorous sense. Musk has put the odds at "10% and rising," which is his own estimate rather than a measured result, and many AI researchers are openly skeptical. Expect a very capable model, not artificial general intelligence. ### Will Grok 5 be free? Probably not for full access at launch. Grok keeps a free, rate-limited tier on its older models, and Grok 5 will most likely sit behind paid plans like SuperGrok at first. xAI has not announced Grok 5 pricing yet. ## Sources - Grok (chatbot) - Wikipedia - Grok version history and current model status - `en.wikipedia.org/wiki/Grok_(chatbot)` - xAI Model Documentation - official current models, context windows, and API pricing - `docs.x.ai/docs/models` - Teslarati: Musk's 10% AGI estimate for Grok 5 - the AGI claim, October 2025 - `teslarati.com/elon-musk-grok-5-now-has-10-percent-chance-of-becoming-worlds-first-agi` - Futurism: OpenAI researcher on Musk's AGI claim - expert skepticism - `futurism.com/artificial-intelligence/openai-researcher-mocks-elon-musks-agi` - TechCrunch: SpaceX acquires xAI - the February 2026 merger - `techcrunch.com/2026/02/02/elon-musk-spacex-acquires-xai-data-centers-space-merger` --- ## Grok vs ChatGPT: Which Is Better? Honest Comparison (August 2026) > Grok vs ChatGPT, compared honestly and current to August 2026: real-time data, coding, writing, context, safety, pricing, and which AI to optimize your brand for. - Canonical: https://geotoolbox.ai/blog/grok-vs-chatgpt - Published: 2026-06-24 · Updated: 2026-08-13 Grok vs ChatGPT is usually framed as a fight with a winner. It is not one. As of August 2026, OpenAI's ChatGPT is the polished all-rounder with the deepest ecosystem, while xAI's Grok is the fast, X-native one with real-time data and a far looser filter. They are built for different jobs, and the right pick depends entirely on what you spend your day doing. This is the comparison done straight, with the one question that actually affects your marketing: which of the two cites your brand. It is written for people who publish and market, not for people who build models.
![Scorecard showing where Grok and ChatGPT each lean across eight everyday jobs.](/blog/grok-vs-chatgpt/grok-vs-chatgpt-scorecard.png)
The tale of the tape, August 2026: each assistant leans on different jobs, and most heavy users end up running both.
## Grok vs ChatGPT at a Glance Here is the short version before the detail. Both are strong, the real differences sit at the edges, and model versions move monthly, so check the update date at the top of this page before you act on anything below. Two launches landed in early July: OpenAI's GPT-5.6 (Sol) became generally available on July 9, 2026, and now leads the ChatGPT side, superseding GPT-5.5, while xAI shipped [Grok 4.5](https://geotoolbox.ai/blog/grok-4-5) the day before, on July 8 (EU access followed on July 16, after an initial AI Act delay).
 ChatGPTGrok
MakerOpenAIxAI (now part of SpaceX)
Top model (August 2026)GPT-5.6 (Sol, Terra, Luna)Grok 4.6
Leans best atCoding, polished writing, a deep tool and integration ecosystemReal-time data from X, speed, cheap high-volume output, native image and video
Real-time dataWeb search through ChatGPT Search when invokedNative X feed plus DeepSearch when enabled
Context windowAbout 1M tokens on GPT-5.6About 500K tokens on Grok 4.6
Safety and toneStrict, cautious, more likely to refuseLighter filter and a looser default tone, with a documented incident history
Image and videoGenerates images natively (GPT Image)Generates images (Aurora) and video (Grok Imagine)
Entry priceGo around $8; Plus around $20/monthFrom around $8 (X Premium) or $10 (SuperGrok Lite); about $30 for SuperGrok
Free tierYes (a GPT-5-class model with limits)Yes (limited usage)
Both are [large language models](https://geotoolbox.ai/glossary/large-language-model) built on the same basic idea, so a feature checklist only gets you so far. The differences that decide which one you will prefer come from how each company thinks and how each handles the live web. One thing to flag up front: most "Grok vs ChatGPT" comparisons you will find run on stale versions or a months-old spec table. Everything here is kept current, which by itself changes several of the usual answers. ## The Makers: OpenAI and xAI Want Different Things The two companies were built around opposite instincts, and that is the root of almost every difference you will feel. ChatGPT comes from OpenAI, the lab that started the consumer AI boom and has spent the years since building scale, polish, and a wide developer ecosystem around it. Its product instinct is to be the reliable default for the most people, which shows up as caution and consistency. Grok comes from xAI, the company Elon Musk founded in 2023 with the stated goal of a "maximally truth-seeking" AI that says things other assistants will not. The history sharpens the rivalry. Musk co-founded OpenAI in 2015, left after a falling-out, and built xAI partly as a direct answer to it. So this is not two neutral labs converging on the same product. It is two opposing bets: OpenAI optimizing for a trusted, broadly useful assistant, and xAI optimizing for one that is fast, plugged into the live conversation, and willing to answer. The corporate picture behind Grok also moved fast, with xAI folded into SpaceX as its AI division during 2026, making Grok's tie to the X social network structural rather than a bolt-on. That contrast runs through everything below. If you want the deeper background on each, we cover [Grok](https://geotoolbox.ai/blog/what-is-grok) and [how ChatGPT actually works](https://geotoolbox.ai/blog/how-does-chatgpt-work) in their own guides. ## Real-Time Data: Grok's Structural Edge, and Where It Breaks **This is Grok's signature advantage, and it is worth separating into two things comparisons usually blur together.** The first is plain web search, which both assistants have: ask either about something recent and it can fetch pages and summarize them through ChatGPT Search or Grok's web access. The second is a native, real-time connection to X, the social network xAI owns, plus its DeepSearch agent that scans X and the open web. That live social feed is the part ChatGPT cannot match, because OpenAI does not own a social platform. For some jobs that gap is decisive. To read the mood on a breaking story, track a launch as it happens, or pull live sentiment from a trend, Grok has no real equivalent. But there is a caveat the hype skips. Real-time access is not the same as reliable research, and testing has caught Grok out. In hands-on testing, Grok has returned news over a week old while ChatGPT surfaced stories from the past couple of days with more reputable sources. A live social feed is fast and noisy, full of unverified claims, and Grok's retrieval depends on the mode and tools in play rather than firing on every query. Live data is excellent for "what is being said" and shakier for "what is true." For the mechanics of how either assistant pulls and ranks live pages, see [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work). Either way, treat Grok as your window into the live conversation, and verify what it finds before you publish. ## Reasoning and Coding **On coding the two trade benchmark wins, but hands-on reports and user feedback tilt real-world reliability toward ChatGPT.** OpenAI has leaned hard into developer work: ChatGPT pairs the chat model with Codex, a dedicated coding agent, and in hands-on tests its code tends to run clean the first time. In hands-on testing, ChatGPT has produced a working tool with no edits while Grok's version shipped a broken feature, even if Grok's built-in error detection then fixed it. One test is not proof, but it matches the pattern. Grok is a genuinely capable frontier reasoning model, and on public benchmarks the two are close enough that the lead flips from test to test. Its real advantages are speed and price: a much cheaper API makes high-volume generation affordable, and for a quick script or a fast iteration loop it is responsive. The recurring complaint among Grok users is that its real-world code misses edge cases, skips parts of the spec, and needs more cleanup than the benchmarks imply, which is why many developers, including parts of Grok's own community, still route serious work to ChatGPT or Claude. So the durable rule is about the shape of the work, not a leaderboard. Reach for ChatGPT when correctness on a multi-step problem is the point, and Grok when speed and volume matter more than a polished first pass. ## Writing and Tone The writing difference is the cleanest split in the comparison. ChatGPT tends to be the stronger all-round writer: it holds structure across long pieces, follows instructions closely, and now exposes explicit tone controls, so a report, a nuanced email, or documentation usually needs less editing. Grok writes with a different voice on purpose: punchy, conversational, often unfiltered, and tuned for the rhythm of social posts. For a sharp take on a trend or copy that should sound like a person rather than a brand, that voice is an asset. There is a nuance worth correcting, because people repeat the opposite. Grok "feels" more human, so it gets called the better creative writer, yet measured quality is more contested than that reputation: ChatGPT leads Grok on some public creative-writing leaderboards while Grok leads on others. Voice and measured quality are not the same thing. The other catch is brand safety: the voice that works in a personal feed can read as a liability under a company name, so anything Grok writes for publication needs a closer read before it ships. ## Context Window and Limits This is where the July 2026 launches flip the usual answer. Through mid-2026 the honest answer was effectively a tie: both flagships sat near a million tokens. The recent Grok updates reopened a gap in ChatGPT's favor, because Grok 4.6 ships a smaller window. Per [xAI's own model docs](https://docs.x.ai/developers/models), [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) (the current flagship while [Grok 5](https://geotoolbox.ai/blog/grok-5) trains) handles up to 500K tokens, while [GPT-5.6's published window](https://developers.openai.com/api/docs/models/gpt-5.6-sol) is **1,050,000** tokens, with 128,000 max output. That is not an estimate: OpenAI states it in the model documentation for all three variants. At the flagship tier ChatGPT now holds a window more than twice the size, though both are big enough for the jobs most people throw at them. A million tokens is enough to load an entire codebase or a stack of reports into one conversation. Two practical notes, though. These are the API ceilings, and the consumer chat apps often expose less and throttle context by plan. And a big window is not the same as flawless recall. In our experience auditing how these tools handle long inputs, the number that matters is rarely the headline size; it is how cleanly each model uses what you actually give it. ## Safety, Guardrails, and Brand Risk **The reason ChatGPT feels cautious and Grok feels loose is two different design choices, and one of them carries real reputational risk.** OpenAI tunes ChatGPT toward strict, institutional safety: it is predictable in a classroom, a clinic, or a corporate blog, and the cost is that it sometimes over-refuses or over-warns on perfectly benign requests, a recurring gripe from its own users. Grok was built around a "truth-seeking" goal and a deliberately lighter filter, so it answers more directly, including things other assistants decline. That openness is the selling point and the hazard. Grok has a documented string of content-moderation failures recorded on its [Wikipedia page](https://en.wikipedia.org/wiki/Grok_(chatbot)): in May 2025 it injected "white genocide" claims into unrelated answers, and in July 2025 it produced antisemitic content and praised Hitler after a system-prompt change, which xAI then rolled back. More serious is the [sexual-deepfake scandal](https://en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal) around [Grok Imagine](https://geotoolbox.ai/blog/grok-imagine), whose loosely-filtered image mode was used to generate non-consensual sexual images at scale. xAI restricted the feature to paid subscribers by March 2026 amid mounting lawsuits, and a Dutch court went on to order six-figure daily fines for continued violations, yet reporting through mid-2026 found the problem had not fully stopped. One side effect cuts against Grok's pitch: the scandal forced xAI to tighten its image tools, so the "anything goes" reputation is more complicated than it once was. None of this makes Grok unusable, but publishing its output under your name without review is a measurable brand risk in a way that is less true of ChatGPT, and no assistant is immune from confidently stating something false, the mechanism we break down in [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations). ## Privacy and Your Data If you put client work or anything sensitive into these tools, the data question matters more than any benchmark, and it rarely gets a mention. Both companies let you control whether your conversations are used to train future models, but the defaults and the plumbing differ, so this is worth checking rather than assuming. Grok is wired into your X account, which is the source of its real-time edge and also the thing to watch. X has a setting governing whether your activity and Grok interactions help train the model, and it has tended to default on, so you have to turn it off if you do not want your prompts used. Anything you post publicly on X can be used to train Grok regardless of whether you ever open it. OpenAI offers a comparable control over whether your ChatGPT chats improve its models, and its business, enterprise, and API tiers are not used for training by default, which is one reason ChatGPT shows up more often in regulated settings. Treat both consumer apps as places where your data may train the model unless you have changed the setting, and verify the current policy before trusting either with confidential material. For client deliverables, the enterprise or API tier with training switched off is the safer path. ## Image and Video Generation Both generate visuals, which is one place neither is the obvious winner. ChatGPT creates images natively through its built-in GPT Image generator, strong on instruction-following and text rendering. Grok counters with two tools of its own: Aurora for images and Grok Imagine for short video, so it can produce a clip in the same window where it writes, which ChatGPT's built-in tools do not match as cleanly. The quality split is real, though. In hands-on tests Grok's images skewed toward unintended realism, producing literal, slightly off results when the prompt called for something stylized or cute, and the same loose-filter history that powers Grok Imagine is exactly what landed it in the deepfake scandal above. For most publishing work, dedicated design tools still beat either chatbot's built-in generator, so treat this as a convenience feature rather than a deciding factor. ## Pricing and Plans Sticker prices are close at the entry level, and the real difference shows up on the API. Here is the current consumer picture, dated July 2026 and worth confirming because both companies move these numbers around.
 ChatGPTGrok
FreeYes, a GPT-5-class model with usage limitsYes, limited usage
Entry paidGo around $8; Plus around $20/monthX Premium around $8 or SuperGrok Lite around $10; SuperGrok at about $30
Higher tierPro at $100, up to $200/monthX Premium+ around $40; SuperGrok Heavy around $300
API, per 1M tokensGPT-5.6 (Sol) at about $5 in / $30 outGrok 4.6 at about $2 in / $6 out
On the API the gap is large and runs in Grok's favor. Per [xAI's pricing](https://docs.x.ai/developers/models), Grok 4.6 is roughly $2 per million input tokens and $6 per million output, undercutting [GPT-5.6 (Sol)](https://openai.com/index/previewing-gpt-5-6-sol/) at near $5 and $30. Those Grok rates cover prompts under 200K tokens; past that they double to $4 / $12, which still lands under Sol. If you generate at volume, that difference compounds fast, though caching discounts and how many reasoning tokens each model burns can narrow the real-world gap. The consumer story is more layered. [ChatGPT's standalone Plus](https://chatgpt.com/pricing/) at about $20 undercuts SuperGrok at about $30, so for a single subscription ChatGPT is usually the cheaper full-featured tier. But Grok has cheaper ways in if you already live on X: a Premium bundle around $8 and a SuperGrok Lite tier around $10, both under Plus. The right way to think about it is a break-even. Grok's cheaper tokens and entry tiers win when you need a lot of decent output, and ChatGPT's reliability earns its premium when an error is expensive. Price the task, not the subscription, and be skeptical of the $200 to $300 top tiers unless you can name exactly what extra you are buying. For the full Grok breakdown, see our [Grok pricing guide](https://geotoolbox.ai/blog/grok-pricing). ## Which Should You Use? **Skip "it depends." The choice is predictable once you name the job.** It comes down to what you spend most of your time doing and how much an error costs you. Choose **ChatGPT** if your work is: - Coding where the first pass needs to run, backed by a mature tool ecosystem - Polished, structured writing and documents that have to be brand-safe - Anything that leans on integrations, custom assistants, or file and data analysis - General-purpose work where reliability beats novelty Choose **Grok** if your work is: - Tracking live discussion, breaking news, or sentiment on X right now - Fast, high-volume drafting where cheap tokens beat a perfect answer - Generating images or short video in the same place you write - Quick prototyping where speed matters more than polish And genuinely consider **both**, which is what most heavy users land on: ChatGPT for the careful, publishable work, Grok for the real-time read and the cheap bulk drafting. If you want the same straight treatment of the other big assistants, our [Grok vs Gemini](https://geotoolbox.ai/blog/grok-vs-gemini), [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude), [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt), and [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) breakdowns line up beside this one. ## What Grok vs ChatGPT Means for Your Brand's Visibility **Here is the question every other comparison skips, and the one that matters most if you market or publish: when someone asks Grok or ChatGPT about your industry, which one mentions you, and why.** The two decide what to cite in almost opposite ways, so being visible in one tells you very little about the other. Neither company publishes its citation logic, so treat what follows as informed analysis, not vendor fact. ChatGPT, through ChatGPT Search, leans on the indexed web and tends to surface sources that read as authoritative and well-referenced, with links, much like a careful search engine. Its reach is enormous, so a mention there is seen by a lot of people. AI answers on a query like this one tend to lean on Wikipedia, official pages, and established tech press, the kind of source that bar rewards. Grok leans the other way, weighting its native X connection: social signals, what is being posted and shared right now, and the live web, which means an active, talked-about presence on X gives you a real shot at being surfaced. One playbook will not win you both, and the same split shows up when you compare [how ChatGPT and Perplexity cite sources](https://geotoolbox.ai/blog/chatgpt-vs-perplexity). The table below reflects those observed tendencies, not published ranking rules.
 To show up in ChatGPTTo show up in Grok
What it appears to weightAuthoritative, well-referenced pages; the indexed web via ChatGPT SearchX and social signals, trending posts, the live web
What earns a mentionClear, factual, citable pages a cautious model trustsReal-time relevance and social proof on X
Where to investAuthoritative coverage, consistent facts, reference-grade pagesAn active, discussed presence on X
This is the work geotoolbox was built for. It tracks whether AI engines, ChatGPT and Grok included, are citing your brand, so you can see the gap instead of guessing at it. If you are starting from scratch, our guides on [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo), [showing up in ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt), and [measuring your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) lay out the playbook for each engine. ## Frequently Asked Questions ### Why do people use Grok instead of ChatGPT? Mostly for two things: real-time data and a looser filter. Grok's native connection to X makes it strong for live news, trends, and social sentiment, and its lighter guardrails mean it answers things ChatGPT declines. It is also cheaper to use at high volume on the API. For everyday all-round work, most people still find ChatGPT the more reliable default. ### What is more reliable, ChatGPT or Grok? ChatGPT, in most hands-on comparisons and user reports. It tends to follow instructions more closely, code more cleanly, and stay grounded on facts, while Grok is faster and more opinionated but more prone to confident errors outside live topics. For anything where being wrong is costly, ChatGPT is the safer pick. ### How much is Grok per month, and is it free? Grok has a free tier, but it is tight, with only a small number of prompts per few-hour window, so most people who use it seriously end up paying. Paid access starts around $8 through X Premium or about $10 for SuperGrok Lite, with the fuller SuperGrok tier near $30 a month and a SuperGrok Heavy tier at about $300. ChatGPT's free tier is more generous and the usual way people try it, and its comparable Plus plan is around $20. ### What can Grok do that ChatGPT can't? Pull live data straight from X, the social platform xAI owns, which ChatGPT has no equivalent for. Grok also generates short video through Grok Imagine in the same window where it chats. ChatGPT can search the web and generate images, but it has no native social-media feed. ### Is ChatGPT or Grok better for coding? ChatGPT for most real work. It pairs with the Codex agent, tends to produce code that runs the first time, and handles multi-step problems more reliably. Grok is fast and cheap and fine for quick scripts, but its real-world coding is a recurring complaint among its users. ### Which should I optimize my brand to show up in, Grok or ChatGPT? Both, with different tactics. ChatGPT appears to reward authoritative, well-referenced pages, while Grok seems to weight an active presence and social proof on X. Neither company publishes its ranking rules, so treat this as informed analysis and track your presence in each engine separately, since being cited by one does not predict the other. ## The Bottom Line Grok vs ChatGPT is not a contest with a winner; it is a fork. ChatGPT is the polished, integrated, reliable all-rounder you reach for when the work has to ship clean; Grok is the fast, plugged-in one you reach for when the live conversation and cheap volume matter more than a perfect answer. Match the tool to the job and most people end up using both. There is one more thing you cannot afford to guess at: what these engines say about you. The same brand can be cited confidently by one and invisible in the other, and you cannot fix a gap you cannot see. geotoolbox [tracks your brand's AI citations](https://geotoolbox.ai/features/citation-interceptor) across engines, ChatGPT and Grok included, so you know exactly where you show up and where you need to earn your way in. ## Sources - xAI - Grok models and API pricing - `docs.x.ai/developers/models` - OpenAI - Introducing GPT-5.5 - `openai.com/index/introducing-gpt-5-5` - ChatGPT - plans and pricing - `chatgpt.com/pricing` - Wikipedia - Grok (chatbot): model lineup, real-time features, and incident history - `en.wikipedia.org/wiki/Grok_(chatbot)` - Wikipedia - Grok sexual deepfake scandal - `en.wikipedia.org/wiki/Grok_sexual_deepfake_scandal` --- ## What Is x402? Coinbase's Payment Protocol for AI Agents > x402 lets AI agents pay in stablecoins over HTTP. What it is, how the 402 handshake works, whether it's a token, and how it compares to ACP, AP2 and MCP. - Canonical: https://geotoolbox.ai/blog/what-is-x402 - Published: 2026-06-24 · Updated: 2026-07-20 If you have seen "x402" next to AI agents and stablecoins and could not tell whether it was a protocol, a product, or a coin, here is the plain answer. x402 is a protocol, and a genuinely new one: a way for software to pay for things over the web, per request, with no human in the loop. This covers what it is, what it is not, and whether it belongs on your radar. ## What Is x402? x402 is an open protocol that lets software pay for things over the web inside a normal web request. It does this by reusing HTTP status code 402, "Payment Required," a slot that was written into the HTTP spec back in the 1990s and then left unused for nearly 30 years. When a client asks for a paid resource, the server answers with a 402 and a price. The client pays, and the server returns the resource. Payment settles in stablecoins, almost always USDC. [Coinbase built x402](https://docs.cdp.coinbase.com/x402/welcome) and launched it on May 6, 2025. In April 2026 it was contributed to the [Linux Foundation](https://www.linuxfoundation.org/x402foundation), and in July 2026 the [x402 Foundation formally launched](https://www.coindesk.com/tech/2026/07/15/visa-mastercard-and-ripple-join-the-standard-letting-ai-agents-pay-in-stablecoins) under neutral governance with 40 members, including Visa, Mastercard, American Express, Stripe, Adyen, Shopify, Google, AWS, Cloudflare, Ripple, Circle, and the Solana and Stellar foundations. That lineup matters: card networks that are building their own agent-payment rails are also backing x402 as a stablecoin settlement layer, a strong signal it has grown beyond a Coinbase project. The clients most likely to use it are [AI agents](https://geotoolbox.ai/glossary/ai-agent): software that calls APIs, buys data, and shops on a person's behalf and cannot stop to type a card number into a form. One thing to settle up front, because it is the most common mix-up. x402 is a protocol, not a coin. There is no official x402 token. We will come back to why so many people think there is. ## Why x402 Exists: The Web Never Had a Native Payment Layer The web shipped without a way to charge for a single request. The 402 status code was meant to fill that gap and never did, because there was no money that moved at the speed of an HTTP call. So the web routed around it. Sites bolted on accounts, subscriptions, API keys, and card processors, all of which assume a human is present to sign up and approve. That assumption breaks twice over for machines. An agent cannot read an SMS code or fill a checkout form, so card rails stop it at the gate. And the economics never worked for small amounts: when a processor takes around $0.30 to move $0.01, charging a fraction of a cent per API call is absurd. Micropayments stayed a nice idea that the rails could not carry. x402 targets exactly that gap. Cloudflare, whose network sees [over a billion 402 responses a day](https://blog.cloudflare.com/x402/) sent to bots, frames it as giving the web a way for clients and servers to exchange value in a common language. The point is not crypto for its own sake. It is letting one machine pay another, per request, without an account in the middle. ## How x402 Works: The 402 Handshake
![Four-step x402 handshake: request, 402 with terms, signed stablecoin payment, 200 with the resource.](/blog/what-is-x402/x402-payment-handshake.png)
Two round trips: the server names its price in a 402, the agent pays, and the resource comes back.
The whole flow is a quick back-and-forth, two round trips with a payment in between. Here is the plain version, per [Coinbase's documentation](https://docs.cdp.coinbase.com/x402/core-concepts/how-it-works): 1. An agent requests a resource, the same as any API call 2. The server replies with **402 Payment Required** and the terms: the price, which network and asset to use, and where to send it 3. The agent signs a stablecoin payment and sends the request again, this time with the payment attached 4. A facilitator checks the payment and settles it on chain, and the server returns a normal 200 response with the resource No account, no API key, no redirect to a checkout page. On fast chains, payments confirm in a few hundred milliseconds, though the exact timing depends on the chain and facilitator. The **facilitator** is worth naming, because it does the heavy lifting. It verifies the payment (the agent signs it with its own wallet key, so a server cannot forge one) and settles it on chain, so the seller does not run blockchain infrastructure. It is technically optional, since a seller can settle on chain itself, but in practice most use a hosted one from Coinbase, thirdweb, or another provider. Coinbase's is free up to 1,000 settlements a month, then $0.001 each. That reliance is also where a lot of the honest criticism lands, which we get to below. The version most explainers describe is v1. The [v2 update](https://www.x402.org/writing/x402-v2-launch) shipped in December 2025 and tidied the mechanics: standardized headers (`PAYMENT-REQUIRED`, `PAYMENT-SIGNATURE`, `PAYMENT-RESPONSE`) instead of the older `X-PAYMENT` style, network identifiers in a common format, and wallet-based sessions so an agent does not re-sign a full payment on every single call. It stays backward compatible with v1. Set against a normal card checkout, the tradeoff is clean but two-sided:
StepCard / Stripe checkoutx402
SetupAccount, KYC, a form to fillOne request, no account
SettlementOften a day or twoRoughly 200 to 400 ms
FeesPercentage plus a fixed fee (around $0.30)No protocol fee; facilitator and network fees apply
RefundsBuilt-in chargebacksNone natively (see caveats)
Built for agentsNo, assumes a humanYes, machine to machine
## Is x402 a Token? No, and Here's the Confusion The protocol has no native token. It settles in stablecoins that already exist, mainly USDC, and it does not need or issue a coin of its own. Coinbase says this plainly in the docs. So what are the "X402" tickers people trade? Unaffiliated. There are speculative tokens floating around on decentralized exchanges that borrow the name, and they have nothing to do with the standard. Buying one does not buy you a piece of the protocol. The confusion got a boost from the numbers. A memecoin called PING on Base could be minted by making an x402 payment, and because minting cost almost nothing, people did it on a loop. That farming inflated x402's early transaction counts and sent the chart near vertical, part of how the protocol later reported over 100 million cumulative payments, most of them on Base. Once the frenzy cooled, activity fell off hard. A large share of that early spike was speculative loop-minting rather than real demand. None of this means x402 is fake. It means the headline metrics are noisy, and you should read any transaction or volume figure with the scope attached. The protocol is real and shipping. The trading narrative wrapped around it is mostly separate. ## What x402 Is Actually Used For Strip away the speculation and the live use cases are narrow but real. They cluster around one idea: charging per call instead of per subscription. A few that exist today. Neynar lets agents pay for individual Farcaster social-data queries. Hyperbolic sells GPU inference by the millisecond, billed through x402. Token Metrics swapped a monthly plan for pay-per-call access to its crypto data. And Cloudflare's Pay Per Crawl uses the same plumbing to let a site charge [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) for access instead of just blocking them. People can use it too, for things like paying a few cents to read one article without a subscription. On the rails: x402 runs on Base by default, with Solana, Polygon, Arbitrum, and others supported. USDC is the practical default. One honest caveat the marketing skips: x402 leans on a signing standard called EIP-3009 (it lets a wallet authorize a token transfer with one signature) that, in practice, only a couple of stablecoins like USDC and EURC support natively. "Works with any token" is true on paper and bumpier in reality, since other assets need extra plumbing. If you see x402 in production, assume USDC on Base until told otherwise. ## x402 vs ACP, AP2, UCP, and MCP This is where most coverage gets muddled, because these names get listed as rivals when they mostly are not. They sit at different layers of the same stack. x402 is the settlement rail. The others handle the steps around payment: finding tools, proving consent, completing a checkout. The clearest example is Google's [Agent Payments Protocol (AP2)](https://ap2-protocol.org/). AP2 handles authorization through signed "mandates" that prove a human told the agent to spend. It does not move money. To actually settle in stablecoins, AP2 can drop down to x402: the [A2A x402 extension](https://github.com/google-agentic-commerce/a2a-x402), built by Google, Coinbase, the Ethereum Foundation, and MetaMask, lets an AP2 flow settle a payment through x402. They stack, they do not compete. The others fit the same way. MCP (Anthropic) connects agents to tools and data and is not a payment system at all. The Agentic Commerce Protocol (OpenAI and Stripe) and Universal Commerce Protocol (Google) handle merchant checkout inside chat surfaces, over existing payment rails. x402 is the layer any of them can drop down to when the payment needs to be a stablecoin moving machine to machine.
ProtocolWhoJobRelation to x402
x402CoinbaseStablecoin payment over HTTPThe settlement rail itself
AP2GoogleAuthorization (signed mandates)Uses x402 for stablecoin settlement
ACP (Agentic Commerce)OpenAI + StripeMerchant checkout, existing railsDifferent layer, can coexist
UCPGoogleEnd-to-end buying in GeminiDifferent layer, can coexist
MCPAnthropicConnect agents to tools and dataNot payments, pairs with x402
One acronym, two protocols. "ACP" usually means the **Agentic Commerce Protocol** from OpenAI and Stripe, above. It can also mean the **Agent Commerce Protocol** from Virtuals, a separate on-chain system for agents that hire each other, using escrow and an evaluator to release payment once work is verified. Neither is x402. x402 is the rail underneath. ## The Honest Caveats: Refunds, Overspend, and Control x402 is genuinely useful and genuinely early, and the gaps matter more than the hype. **No refunds.** On-chain payments are final. There is no chargeback and no central party to reverse a mistake. The protocol has no native dispute mechanism, so refunds depend on bolt-on escrow extensions like x402r rather than anything built in. If you pay and the service errors out or returns junk, nothing claws the money back, which is part of why pay-per-call is used for small, low-stakes amounts today, not big-ticket buys. **Nothing stops an agent overspending.** Spending limits live in the agent or wallet, not in the protocol. x402 will happily process every 402 it is handed, so a buggy loop or a hostile endpoint that keeps returning 402 can drain a budget unless you cap it yourself. The decentralization is also thinner than the pitch suggests. Most deployments lean on a hosted facilitator, usually Coinbase's, which is a convenient single point of trust and a potential chokepoint, the opposite of the permissionless promise. There is a privacy and compliance gap on top. Tying payments to HTTP requests links IP addresses and timestamps to on-chain activity, and x402 itself does no KYC or sanctions screening, so that falls to the facilitator or seller, with the legal and tax liability landing on you. A [security paper](https://arxiv.org/abs/2605.30998) has also documented logic flaws in x402 implementations, from reusing one payment across different resources to race conditions that let a request slip through unpaid. None of this is disqualifying. It is the difference between a working primitive and a finished product. ## What x402 Means for Your Site Most site owners will not touch x402 directly for a while, and that is fine. The useful takeaway is upstream of payment. x402 handles the last step, the paying. It assumes the agent already found you, reached your pages, and understood what you sell. Cloudflare's Pay Per Crawl is the clearest example of why that order matters: it turns the AI crawler question from "block or allow" into "charge," but a crawler can only pay you if it can [reach and read your pages](https://geotoolbox.ai/blog/agent-ready-website) in the first place. A bot that hits a login wall, a CAPTCHA, or a blank JavaScript shell never gets to the 402. In our experience, the brands worth worrying about agent payments are the ones already winning the earlier gates, where an agent can find them, read them cleanly, and trust them enough to recommend them. That is the same groundwork behind [agentic commerce](https://geotoolbox.ai/blog/agentic-commerce) generally. Payment is the easy part to add later. Being legible to a machine is the part you build now. ## The Short Version x402 is a real payment rail with real gaps: a clever revival of HTTP 402 that lets agents pay in stablecoins per request, wrapped in a noisy token narrative it does not actually have. Worth understanding, not worth panicking over. The move that pays off today is making sure agents can find and parse your site at all, long before any of them tries to pay you. If you are not sure they can, our [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows what an AI bot actually sees when it visits. That is the gate x402 quietly assumes you have already passed. ## Frequently Asked Questions ### Is x402 a cryptocurrency or token? No. x402 is a protocol, not a coin, and it has no native token. It settles payments in existing stablecoins, mainly USDC. Any "X402" token you see trading on an exchange is unaffiliated with the protocol. ### Did Coinbase create x402? Yes. Coinbase built x402 and launched it in May 2025. In April 2026 it was contributed to the Linux Foundation, and in July 2026 the x402 Foundation formally launched with 40 members, including Visa, Mastercard, American Express, Stripe, Google, AWS, and Circle, so it is now a neutral, open standard rather than a Coinbase-owned project. ### How is x402 different from paying with a card? There is no account, no KYC, and no checkout form, so software can pay on its own. Settlement is near instant and works for sub-cent amounts. The trade-off is that payments are final: there are no built-in chargebacks or refunds. ### Is x402 the same as MCP, AP2, or ACP? No, and they are not really rivals. MCP connects agents to tools, AP2 proves a user authorized a payment, and ACP handles merchant checkout. x402 is the settlement rail underneath, and AP2 actually uses it to move stablecoins. ### What stops an AI agent from overspending with x402? Nothing in the protocol itself. Spending limits are set in the agent or its wallet, not in x402. If you let an agent pay, you have to cap how much and how often on your side. ### Can humans use x402 or only AI agents? Both. It is designed for machine-to-machine payments, but a person can use it too, for example paying a few cents to read a single article instead of buying a subscription. Either way the payer needs a wallet holding USDC, and for agents that is usually a programmatic server wallet from a provider like Coinbase or thirdweb, not a browser extension. ## Sources - x402 documentation: Overview and How It Works - Coinbase Developer Docs - `docs.cdp.coinbase.com/x402/welcome` - Coinbase debuts x402 for internet-native stablecoin payments - PYMNTS (launch, May 2025) - `pymnts.com/cryptocurrency/2025/coinbase-debuts-x402-internet-native-stablecoin-payments` - Introducing x402 V2 - x402 (open standard) - `x402.org/writing/x402-v2-launch` - Launching the x402 Foundation, and support for x402 transactions - Cloudflare - `blog.cloudflare.com/x402` - The x402 Foundation - Linux Foundation - `linuxfoundation.org/x402foundation` - Visa, Mastercard and Ripple join the standard letting AI agents pay in stablecoins (x402 Foundation formal launch, 40 members) - CoinDesk, July 15, 2026 - `coindesk.com/tech/2026/07/15/visa-mastercard-and-ripple-join-the-standard-letting-ai-agents-pay-in-stablecoins` - Agent Payments Protocol (AP2) and the A2A x402 extension - Google and Coinbase - `ap2-protocol.org` - `github.com/google-agentic-commerce/a2a-x402` - Free-Riding in the AI Economy: Demystifying Logic Flaws in x402-Enabled Payment Systems - arXiv - `arxiv.org/abs/2605.30998` - What is x402? - Ledger Academy - `ledger.com/academy/topics/economics-and-regulation/what-is-x402` --- ## Grok vs Claude: Which Is Better? An Honest Comparison (August 2026) > Grok vs Claude, compared honestly and current to August 2026: real-time data, coding, writing, context, safety, pricing, and which AI to optimize your brand for. - Canonical: https://geotoolbox.ai/blog/grok-vs-claude - Published: 2026-06-22 · Updated: 2026-08-14 Grok vs Claude is usually framed as a fight with a winner. It is not one. As of July 2026, xAI's Grok and Anthropic's Claude are both strong, and they are built for different jobs: Claude leans toward careful reasoning, long documents, and a cautious safety posture, while Grok leans toward real-time data from X, speed, and a far looser filter. This is the comparison done straight, kept current, with one question almost no other comparison asks: which of the two should cite your brand. It is written for people who publish and market, not for people who build models.
![Scorecard showing where Claude and Grok each lean, from context to price.](/blog/grok-vs-claude/grok-vs-claude-scorecard.png)
No overall winner — each model leans ahead on different jobs, and coding benchmarks are a tie.
## Grok vs Claude at a Glance Here is the honest version before the detail. Both are excellent, and the real differences sit at the edges. Model versions move monthly, so check the update date at the top of this page before you act on anything below.
 ClaudeGrok
MakerAnthropicxAI (now part of SpaceX)
Top models (August 2026)Opus 5, Sonnet 5, Haiku 4.5, plus the higher-tier Fable 5 (Opus 4.8 now legacy)Grok 4.6, xAI's current flagship (Grok 4.5 the prior gen)
Leans best atCareful reasoning, long documents, agentic codingFrontier reasoning, real-time data from X, speed, native image and video generation
Real-time dataWeb search when invoked; otherwise training to about May 2026 on Opus 5X Search and DeepSearch when enabled; otherwise its training data
Context window1M tokens on Opus 5 and Sonnet 5500K tokens on Grok 4.5/4.6 (Grok 4.3, still available, is 1M)
Safety and toneConstitutional AI; cautious, more likely to refuseTruth-seeking by design, less filtered, with a documented incident history
Image and videoReads images, does not generate themGenerates images (Aurora) and video (Grok Imagine)
Entry pricePro around $20/monthFrom around $8 (X Premium) or $10 (SuperGrok Lite); about $30 for SuperGrok
Free tierYes (a Sonnet-class model)Yes (limited usage)
Both are [large language models](https://geotoolbox.ai/glossary/large-language-model) built on the same basic idea, so a feature checklist only gets you so far. The differences that predict which one you will prefer come from how each was trained and how each handles the live web. One more thing to flag up front: most "Grok vs Claude" comparisons you will find pair stale versions (Claude Opus 4.6 against Grok 4.1) or a months-old spec table. Everything here is kept current, which by itself changes several of the usual answers. The head-to-head puts [Grok 4.5](https://geotoolbox.ai/blog/grok-4-5) (released July 8, 2026 and pitched by Elon Musk as "Opus-class," in his words "roughly comparable to Opus 4.7, but much faster") against Anthropic's Opus tier. xAI has since shipped [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) (August 12, 2026) as its current flagship, but the head-to-head numbers below were measured on Grok 4.5; Anthropic's Fable 5 tier sits above Opus but does not change the comparison. One timing note: [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5) replaced Opus 4.8 on July 24, 2026 at the same $5 and $25 per million tokens, after Grok 4.5 shipped. Where a benchmark below names Opus 4.8, that is the model the published head-to-head numbers were actually run against, so we have left it as measured rather than relabeling it. ## The Makers: xAI and Anthropic Want Different Things **The two companies were built around opposite instincts, and that is the root of almost every difference you will feel.** [Claude](https://geotoolbox.ai/blog/what-is-claude-ai) comes from Anthropic, an AI-safety company founded in 2021 by former OpenAI researchers. Its whole pitch is careful, steerable models, and that shows up as caution in the product. [Grok](https://geotoolbox.ai/blog/what-is-grok) comes from xAI, the company Elon Musk founded in 2023 with the stated goal of building a "maximally truth-seeking" AI that says things other assistants will not. Same technology, very different north star. The corporate picture behind Grok also shifted fast. xAI acquired X, the former Twitter, in March 2025, and then, [per Wikipedia's record of the deal](https://en.wikipedia.org/wiki/XAI_(company)), SpaceX acquired xAI in an all-stock transaction in February 2026, folding it into SpaceX as its AI division by May 2026. So Grok's maker is now part of the same company that flies rockets, and Grok's tie to the X social network is structural, not a bolt-on. Anthropic, by contrast, has stayed an independent lab focused on one product line. That contrast runs through the rest of this comparison. Anthropic optimizes for a model you can trust with sensitive, long-form work; xAI optimizes for one that is fast, plugged into the live conversation, and willing to answer. Both are valid goals that produce tools which feel different the moment you use them. ## Real-Time Data: Grok's Real Structural Edge **This is Grok's signature advantage, and it is worth separating into two things that comparisons usually blur together.** The first is plain web search, which both assistants have: ask either about something recent and it can fetch pages and summarize them. The second is a native, real-time connection to X, the social network xAI owns, plus its DeepSearch agent that scans X and the open web. That live social feed is the part Claude cannot match, because Anthropic does not own a social platform. For some jobs that gap is decisive. To read the mood on a breaking story, track a launch in real time, or pull live sentiment from a trend, Grok is in a different category. Claude answers from its training, which reaches reliably to about May 2026 on Opus 5, unless you send it to the web. For anything happening right now, that is a real limitation. There is an honest caveat for publishers, though. Real-time access to X is not the same as reliable research. A live social feed is fast and noisy, full of unverified claims, and Grok's retrieval depends on the mode and tools in play rather than firing on every query. Live data is excellent for "what is being said," and shakier for "what is true." Claude's slower, training-grounded answers are often steadier on settled facts. For the mechanics of how either assistant pulls and ranks live pages, see [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work). Either way, treat Grok as your window into the live conversation, and verify what it finds before you publish. ## Reasoning and Coding **On coding the two trade wins, and the right answer depends on whether you mean a benchmark or a real project.** On the public benchmarks the gap between Grok 4.5 and Claude Opus 4.8 is small, and which one is "ahead" flips from test to test: Grok 4.5 leads Opus 4.8 on DeepSWE 1.0 and Terminal Bench 2.1, but trails it on the neutral DeepSWE 1.1 run and on SWE Bench Pro. So this is not a clean win for either side; treat any single leaderboard number as a snapshot. The more durable split shows up in how each behaves on actual work. Claude is the one most teams reach for on complex, multi-step problems: refactoring across many files, reasoning about a large codebase, catching the logical edge case that breaks later. That strength is backed by an agentic ecosystem, Claude Code and its tooling, built for letting the model plan and execute longer tasks. When the job is "understand this whole repository and change it carefully," Claude's depth tends to win. Grok 4.5 is a genuine frontier model, "Opus-class" by xAI's framing and thirteenth of 190 on the Artificial Analysis Intelligence Index as of July 25, 2026, and it was trained in part on Cursor data to sharpen its coding, so this is not a smart-versus-fast story (a full [Grok 5](https://geotoolbox.ai/blog/grok-5) is still in training). It was EU-blocked at launch under the AI Act but became fully available across the EU on July 16, 2026. Its day-to-day edge is speed: Musk pitched it as roughly Opus-class but much faster, and for a quick script, a single-file prototype, or a fast iteration loop it is responsive and gets you a working draft quickly, while its cheaper API makes high-volume generation more affordable. Where it tends to trail is the hard architectural call on a big, messy codebase, though that is a tendency, not a ceiling. So the practical rule is about the shape of the work, not a winner: Claude when correctness on a complicated system is the point, Grok when speed and volume matter more. If you want to understand why Claude's training makes it behave this way, we go deep in [how Claude works](https://geotoolbox.ai/blog/how-does-claude-work). ## Writing and Tone **The writing difference is the cleanest split in the whole comparison.** Claude tends to be the stronger long-form writer: it holds a thread across thousands of words, keeps structure under control, and produces clean prose that needs less editing, though that same caution can make it read a little flat. For a report, a nuanced email, a piece of documentation, or anything where the reader expects a professional register, it is usually the better first draft. Grok writes with a different voice on purpose: punchy, conversational, often unfiltered, and tuned for the rhythm of social posts. For a sharp tweet, a reactive take on a trend, or copy that should sound like a person rather than a brand, that voice is an asset. The catch is brand safety. Grok's looser filter is the same trait that produces the writing voice people like, and it cuts both ways. Output that is fun in a personal feed can be a liability under a company name, so anything Grok writes for publication needs a closer read before it ships. Claude's caution makes it duller in a group chat and safer on a corporate blog. Pick the voice that matches where the words will live. ## Context Window and Output Limits **Being current to August 2026 changes the usual answer here, with a twist.** Older comparisons confidently tell you Claude has a much bigger context window than Grok, or that Grok's is tiny. Neither is quite right. Per [Anthropic's model docs](https://platform.claude.com/docs/en/docs/about-claude/models/overview), Claude Opus 5 and Sonnet 5 each handle up to 1 million tokens. Per [xAI's own model docs](https://docs.x.ai/docs/models), Grok 4.3 matched that at 1 million, though the newer Grok 4.5 and 4.6 flagships ship a 500,000-token window. So Claude's current flagship holds more, but 500K is still enormous, not the tiny window the old framing assumes. A token is roughly three-quarters of a word, so even half a million tokens is on the order of several hundred thousand words, enough to load an entire codebase, a long contract, or a stack of research reports into a single conversation. Both flagship models accept vast inputs on the API, and maximum output runs high too, up to about 128,000 tokens per response on Claude's top models. The old "Grok's window is tiny" framing no longer holds, even if Claude's is now the larger of the two. Two practical notes. First, these are API limits; the consumer chat apps often expose less and throttle context by plan, so what you get in the Grok or Claude app may be smaller than the headline number. Second, a big window is not the same as flawless recall, and both models reason best when the key material sits near the start or end of a long prompt. In our experience auditing how these tools handle long inputs, the gap that matters is rarely raw window size, it is how cleanly each model uses what you give it. ## Safety, Personality, and Brand Risk **The reason Claude feels cautious and Grok feels loose is not vibes, it is two different training recipes, and one of them carries real reputational risk.** Claude is trained with [Constitutional AI](https://www.anthropic.com/news/claudes-constitution), a method where the model first critiques and revises its own answers against a written set of principles, then learns from AI-generated preferences based on those same rules. In Anthropic's words, "the system uses a set of principles to make judgments about outputs, hence the term 'Constitutional.'" That is why Claude is quicker to add a caveat, ask a clarifying question, or refuse on the edges. Grok was built around xAI's "truth-seeking" goal and a deliberately lighter filter, so it answers more things more directly, including things other assistants decline. That openness is the selling point and the hazard. Grok has had a documented string of content-moderation failures, recorded with sources on its [Wikipedia page](https://en.wikipedia.org/wiki/Grok_(chatbot)): in May 2025 it injected "white genocide" claims into unrelated answers, and in July 2025, after a system-prompt change meant to make it less filtered, it produced antisemitic content before xAI rolled the change back. Separately, a misconfiguration in August 2025 let some private user conversations get indexed by Google, and its image tools drew safety and legal scrutiny in late 2025 over nonconsensual sexual imagery. None of that means Grok is unusable, but it does mean publishing its output under your name without review is a measurable brand risk in a way that is less true of Claude. The flip side is real too. Claude's caution becomes friction: over-refusing or over-warning on perfectly benign requests is a recurring complaint from its own users, and it is a real cost when you just want the task done. And no assistant is immune from confidently stating something false, the mechanism we break down in [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations). The honest summary: Claude trades willingness for predictability, Grok trades predictability for openness, and which you want depends on whether the output is for you or for the public. ## Privacy and Your Data **If you put client work or anything sensitive into these tools, the data question matters more than any benchmark, and the two handle it differently.** Both companies let you control whether your conversations are used to train future models, but the defaults and the plumbing are not the same, so this is worth checking rather than assuming. Grok is wired into your X account, which is the source of its real-time edge and also the thing to watch on privacy. X has a setting controlling whether your activity and Grok interactions help train the model, and it has tended to default on, so you have to turn it off if you do not want your prompts used. Anything you post publicly on X is fair game regardless. Anthropic offers a comparable control over whether your Claude chats improve its models, and its business and API tiers are not used for training by default, which is one reason Claude shows up more often in enterprise and regulated settings. Treat both consumer apps as places where your data may train the model unless you have changed the setting, and verify the current policy before trusting either with confidential material. For client deliverables, the enterprise or API tier with training switched off is the safer path. ## Image and Video Generation **Here the competitors flatly contradict each other, so here is the correct version: Grok generates media, Claude does not.** Grok generates its own media: an image model, Aurora, from December 2024, and [Grok Imagine](https://geotoolbox.ai/blog/grok-imagine) for image and video that followed in 2025, both documented on its [Wikipedia page](https://en.wikipedia.org/wiki/Grok_(chatbot)). Inside Grok you can describe a picture or a short clip and get one back. Claude works the other way around. It reads images. Hand it a screenshot, a chart, or a photo and it will analyze what is in it, but it does not generate images or video. That is a deliberate product choice by Anthropic, not a temporary gap, so if your workflow needs a model that produces visuals in the same window where it writes, Grok is the one that does it and Claude is not in the running. For most publishing work this matters less than it looks, because dedicated image and video tools still outclass a chatbot's built-in generator. But if "draft the copy and rough out a visual in one place" is the job, it is a real difference that several other comparisons get backwards. ## Pricing and Plans **Sticker prices are close at the entry level, and the real cost difference shows up on the API.** Here is the current consumer picture, dated July 2026 and worth confirming because both companies move these numbers around.
 ClaudeGrok
FreeYes, a Sonnet-class model with usage limitsYes, limited usage
Entry paidPro, around $20/month ($17 with annual billing)X Premium around $8 or SuperGrok Lite around $10; SuperGrok at about $30 for fuller access
Higher tierMax, from $100/month, up to $200/month for the most usageX Premium+ around $40; SuperGrok Heavy around $300
API, per 1M tokensOpus 5 at $5 in / $25 out; Sonnet 5 at $2 / $10 (permanent)Grok 4.5 at $2 in / $6 out (Grok 4.3, still available, is $1.25 / $2.50)
On the API the gap is still large. Per [xAI's pricing](https://docs.x.ai/docs/models), Grok 4.5 runs $2 per million input tokens and $6 per million output tokens, while per [Anthropic's pricing](https://claude.com/pricing) Claude Opus 5 is $5 and $25, well over double on input and roughly four times on output. Those Grok rates apply to prompts under 200K tokens; above that xAI charges $4 / $12, so the advantage roughly halves on long-context work. So Grok 4.5 gives you Opus-class capability by xAI's billing at well under Opus rates either way: pricier than the prior Grok 4.3, but still far under Opus. If you are generating at volume, that difference compounds fast and Grok is the cheaper engine on sticker rates, though caching discounts and how many reasoning tokens each model burns can narrow the real-world gap. The consumer story is more layered, and in Grok's favor at the bottom. Grok actually has cheaper ways in than Claude: it is bundled into X Premium at about $8 and sells a SuperGrok Lite tier around $10, both under Claude Pro's $20. The fuller SuperGrok tier at about $30 sits just above Claude Pro, with a $300 Heavy tier at the top. So at the entry level Grok can be cheaper, especially if you already pay for X. The honest way to think about it is a break-even: Grok's cheaper tokens and tiers win when you need a lot of decent output, and Claude's reasoning advantage earns its premium when an error is expensive. Price the task, not the subscription. Our [Grok pricing guide](https://geotoolbox.ai/blog/grok-pricing) details every Grok tier. ## Which Should You Use? **Skip "it depends." The choice is predictable once you name the job.** It comes down to what you spend most of your time doing and how much an error costs you. Choose **Claude** if your work is: - Long-document reasoning, drafting, and editing where prose quality matters - Complex, multi-file coding where catching the edge case is the point - Client-facing or regulated output that has to be brand-safe and defensible - Anything you would rather not have to double-check for a confidently wrong answer Choose **Grok** if your work is: - Tracking live discussion, breaking news, or sentiment on X right now - Fast, high-volume drafting where cheap tokens beat a perfect answer - Generating images or short video in the same place you write - Quick prototyping where speed matters more than architectural rigor And genuinely consider **both**, which is what most heavy users land on. A common split is Claude for the careful, publishable work and Grok for the real-time read and the cheap bulk drafting. Both are cheap enough at the entry tier that paying for two often beats the time lost forcing one tool into a job it is bad at. The one real cost of splitting is context: a long working session you build up in one model does not carry over to the other, so keep big, reusable material where you will actually reuse it. If you also want this comparison against the other major assistant, our [Grok vs Gemini](https://geotoolbox.ai/blog/grok-vs-gemini), [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt), and [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) breakdowns use the same honest lens. ## What Grok vs Claude Means for Your Brand's Visibility **Here is the question every other comparison skips, and the one that matters most if you market or publish: when someone asks Grok or Claude about your industry, which one mentions you, and why.** The two decide what to cite in almost opposite ways, so being visible in one tells you very little about the other. Neither company publishes its citation logic, so treat what follows as informed analysis, not vendor fact. Grok leans on its native X connection: it appears to weight social signals, what is being posted and shared right now, and the live web, which means an active, talked-about presence on X gives you a real shot at being surfaced. Claude leans the other way. Its more conservative product behavior and training-grounded answers appear to apply a stricter credibility bar, so it tends to favor sources that read as authoritative and well-established over whatever is trending. That shapes [how to get cited in Claude](https://geotoolbox.ai/blog/claude-seo). In our analysis of where these systems pull answers, the split is clear: socially-weighted AI results surface forums, video, and discussion, while authority-weighted engines lean on reference sites, vendor docs, and established press. One playbook will not win you both. The table below reflects those observed tendencies, not published ranking rules from either company.
 To show up in GrokTo show up in Claude
What it weightsX and social signals, trending posts, the live webAuthoritative, credible, established sources; web search when invoked
What earns a mentionReal-time relevance and social proof on XClear, factual, well-referenced pages a cautious model trusts
Where to investAn active, discussed presence on XAuthoritative coverage, consistent facts, citable reference pages
There is also distribution. Grok's reach is concentrated among X users; Claude reaches through its app, the API, and a large enterprise footprint, so part of the choice is which engine the people you want to reach actually use. This is the work geotoolbox was built for: it tracks whether AI engines, Grok and Claude included, are citing your brand, so you can see the gap instead of guessing at it. If you are starting from scratch, our guides on [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) and [measuring your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) lay out the playbook for each engine. ## Frequently Asked Questions ### Is Grok better than Claude for coding? It is contested and close. On public benchmarks the two trade wins, but for complex, multi-file engineering most teams prefer Claude and its Claude Code tooling, while Grok is strong for fast, single-file prototyping and cheaper high-volume generation. Pick Claude for depth, Grok for speed. ### Is Claude safer than Grok? By design, yes. Claude is trained with Constitutional AI and is more likely to caution or refuse, while Grok runs a deliberately lighter filter and has a documented history of content-moderation failures. For anything published under a company name, Claude is the lower-risk default and Grok output needs a closer human review. ### Which is cheaper, Grok or Claude? On the API, Grok is much cheaper: Grok 4.5 runs $2 / $6 per million tokens versus $5 / $25 for Claude Opus 5, about 76% less on output, with near-Opus results on several benchmarks (those were measured against the prior Opus 4.8). That gap narrows on long prompts, since Grok's rate doubles to $4 / $12 above 200K tokens, which still leaves it around half the price of Opus rather than a quarter. On consumer plans it is mixed: Grok's entry tiers undercut Claude, with X Premium around $8 and SuperGrok Lite around $10 versus Claude Pro at about $20, while fuller Grok access through SuperGrok runs about $30. For light use Grok is cheaper, for heavy API use much cheaper. ### Does Grok have real-time data that Claude does not? Both can search the web, but only Grok can tap a native, real-time connection to X when its search tools are on, its standout advantage for live and trending topics. Claude answers from training, reaching reliably to about May 2026 on Opus 5, unless you send it to the web. For breaking news and social sentiment, Grok has the edge. ### Does Grok train on my data? Does Claude? Both can use your conversations to train future models unless you opt out, and the defaults differ. Grok is tied to your X account, where a setting controlling training has tended to default on, and your public posts are always fair game. Anthropic offers a similar control and excludes business and API traffic by default. Verify the current setting before using either for sensitive work. ### Which AI should I optimize my brand to show up in, Grok or Claude? Both, but with different tactics, because they appear to cite sources differently. Grok seems to reward an active, talked-about presence on X and social signals, while Claude leans toward authoritative, well-referenced pages. Neither company publishes its ranking rules, so treat this as informed analysis and track your presence in each engine separately, since being cited by one does not predict the other. ## The Bottom Line Grok vs Claude is not a contest with a winner, it is a fork. Claude is the careful one, the model you trust with long, sensitive, publishable work; Grok is the fast, plugged-in one, the model you reach for when the live conversation and speed matter more than a perfect answer. Match the tool to the job and most people end up using both. There is one more thing you cannot afford to guess at: what these engines say about you. The same brand can be cited confidently by one and invisible in the other, and you cannot fix a gap you cannot see. geotoolbox [tracks your brand's AI citations](https://geotoolbox.ai/features/citation-interceptor) across engines, Grok and Claude included, so you know exactly where you show up and where you need to earn your way in. ## Sources - Anthropic - Claude models overview (current models, context windows, API pricing) - `platform.claude.com/docs/en/docs/about-claude/models/overview` - Anthropic - Claude plans and pricing - `claude.com/pricing` - Anthropic - Claude's Constitution (Constitutional AI) - `anthropic.com/news/claudes-constitution` - xAI - Grok models and API pricing - `docs.x.ai/docs/models` - Wikipedia - Grok (chatbot): model lineup, real-time features, and incident history - `en.wikipedia.org/wiki/Grok_(chatbot)` - Wikipedia - xAI (company): ownership and the SpaceX acquisition - `en.wikipedia.org/wiki/XAI_(company)` --- ## What Is DeepSeek? China's Open-Weight AI, Explained > What is DeepSeek? The Chinese open-weight AI explained: R1, V3, and V4, the $6M myth, is it safe, the China and bans question, and what it means for you. - Canonical: https://geotoolbox.ai/blog/what-is-deepseek - Published: 2026-06-22 · Updated: 2026-08-23 DeepSeek is the Chinese AI lab whose cheap, open-weight models briefly wiped a record amount off Nvidia's value and made the rest of the industry nervous. If you remember the January 2025 panic and want the plain version of what DeepSeek actually is, who builds it, whether it is safe, and where it stands now, this is it, current as of August 2026, now that [DeepSeek V4](https://geotoolbox.ai/blog/deepseek-v4) has shipped in full. That last part matters. Almost every "what is DeepSeek" article you will find was written during the R1 frenzy in early 2025 and stops there. We will cover the parts they miss or get wrong: the current model lineup, the honest story behind the famous "$6 million" price tag, the [open weights](https://geotoolbox.ai/glossary/open-weights) versus [open source](https://geotoolbox.ai/glossary/open-source-ai) distinction, the real shape of the safety and China questions, and what a strong Chinese open model means for whether AI tools mention your brand. ## What Is DeepSeek? **DeepSeek is a Chinese AI lab in Hangzhou that builds open-weight [large language models](https://geotoolbox.ai/glossary/large-language-model), and the chatbot that runs on them.** Its best-known models are DeepSeek-R1, a reasoning model, and the DeepSeek-V3 and V4 families for general use. The company was founded and is chiefly funded by High-Flyer, a quantitative hedge fund. Like Kimi, DeepSeek is really two things, and keeping them straight clears up most of the confusion. There is the hosted product, the free chatbot at chat.deepseek.com and the paid API, which DeepSeek runs on its own servers. And there are the open-weight models underneath, which DeepSeek publishes so anyone can download, run, and adapt them. The model is open; the service around it is DeepSeek's. That split is the key to almost every question people ask about DeepSeek, because the answers are often different for the hosted app than for the open weights you run yourself. Where DeepSeek earned its reputation is doing frontier-grade reasoning and coding at a fraction of what the closed American models cost, which is exactly why its arrival rattled the market. We break the current numbers down in our [DeepSeek pricing guide](https://geotoolbox.ai/blog/deepseek-pricing). Calling it "the cheap Chinese ChatGPT" undersells both what it did and what is genuinely worth questioning about it. ## Who Makes DeepSeek? High-Flyer and Liang Wenfeng DeepSeek comes from an unlikely parent: a hedge fund. It was founded in July 2023 by [Liang Wenfeng](https://en.wikipedia.org/wiki/DeepSeek), who had already co-founded High-Flyer, a Chinese quantitative fund that traded using AI and had stockpiled Nvidia GPUs for years before US export controls tightened. High-Flyer spun its AI research lab out into DeepSeek, and Liang runs both. The lab is based in Hangzhou and stayed deliberately small, with around 160 staff, hiring researchers fresh out of top Chinese universities rather than expensive veterans. That origin explains a lot about how DeepSeek behaves. It describes itself as research-first and has been in no rush to commercialize, which is part of why it gives so much away for free. The funding has followed the fame: DeepSeek closed its first outside funding round around June 2026 at a reported valuation near $50 billion (about $7.4 billion raised), and within weeks reports pointed to a further round targeting a valuation near $74 billion ahead of a planned IPO. DeepSeek does not confirm figures, so treat any number as reported rather than official. So yes, DeepSeek is a Chinese company, and that fact sits underneath the safety and data questions we get to below. But it is worth being precise about ownership: DeepSeek is owned by High-Flyer and run by Liang, not by the Chinese state, even though, like any Chinese company, it operates under Chinese law. ## The DeepSeek Lineup: R1, V3, and V4 DeepSeek runs two model lines. The V-series (V3, V4) are general-purpose models. The R-series (R1) are reasoning models that work through a problem step by step before answering. The breakout was DeepSeek-R1 in January 2025, which matched OpenAI's o1 on key reasoning and math benchmarks at a tiny fraction of the price. Since then the lineup has moved fast, and this is where stale explainers fall down. Here is where it stands.
Model (as of August 2026)ReleasedWhat it is
DeepSeek-V3December 2024General model: 671B total / 37B active (mixture-of-experts), 128K context
DeepSeek-R1January 2025The reasoning model that started the panic; matched OpenAI o1 on key benchmarks at far lower cost
DeepSeek-V3.1August 2025Hybrid model with both thinking and fast non-thinking modes
DeepSeek-V3.2December 2025More efficient long-context attention
DeepSeek-V4 (generally available since August 13, 2026)April 2026 preview, August 2026 GAV4-Pro (1.6T total / 49B active) and V4-Flash (284B / 13B), both 1M-token context
A few things to know. Most of these are mixture-of-experts models: they hold a huge number of parameters but only switch on a small slice for any given [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), which is the trick behind their low running cost. The big April 2026 jump was the [V4 preview](https://api-docs.deepseek.com/news/news260424), which pushed the context window to one million tokens and split into a fast Flash model and a heavyweight Pro model, both released as open weights under the MIT License. V4 ran as a preview from April 24 and its flagship V4-Pro reached general availability on August 13, 2026, with the Flash model refreshed to the V4-Flash-0731 build at the end of July. (Oddly, the V4-Pro model card still describes "a preview version of DeepSeek-V4 series," even though the product shipped.) On July 24, 2026 the legacy `deepseek-chat` and `deepseek-reasoner` aliases retired. For the full picture of the shipped model, its benchmarks, pricing, and how it compares, see our [DeepSeek V4 guide](https://geotoolbox.ai/blog/deepseek-v4). One model people keep asking about is missing: DeepSeek-R2. It was expected in 2025, but as of mid-2026 it has not shipped. Reporting points to Liang being unsatisfied with its performance and to hardware snags from a push to train on domestic Huawei chips. Because the lineup moves every couple of months, treat this table as an August 2026 snapshot and check the model card for the version you actually use. ## Why DeepSeek Shook the Market, and the $6M Myth
![Scale bars comparing DeepSeek's $5.576M training-run figure with its estimated $1.6B infrastructure.](/blog/what-is-deepseek/deepseek-cost-myth-scale.png)
The famous $6M was one training run; the estimated real infrastructure bill runs past $1.6B.
When the DeepSeek app hit number one on the US App Store in late January 2025, the reaction was not really about the chatbot. It was about cost. DeepSeek claimed it had trained a frontier-class model for a few million dollars, which, if true, undercut the assumption that only companies spending billions could compete. Markets took it literally: Nvidia fell about 17% in a day and lost on the order of $600 billion in value, and commentators called it AI's "Sputnik moment." The efficiency is real. DeepSeek leaned hard on mixture-of-experts design, low-precision math, and clever engineering to train competitive models on the weaker chips it could legally buy. But the famous price tag needs an asterisk. **The "$6 million" figure (more precisely $5.576 million) was the cost of one final training run for V3, not the cost of building DeepSeek.** It excludes the research, the failed runs, the staff, and the hardware. By [SemiAnalysis's estimate](https://www.tomshardware.com/tech-industry/artificial-intelligence/deepseek-might-not-be-as-disruptive-as-claimed-firm-reportedly-has-50-000-nvidia-gpus-and-spent-usd1-6-billion-on-buildouts), DeepSeek sits on around 50,000 Nvidia GPUs and well over a billion dollars in infrastructure. The breakthrough was genuine; the headline number was the smallest true number available, not the real bill. There is a second asterisk on how it got so good so cheaply. DeepSeek's own R1 paper describes training on reasoning data generated by other models, and OpenAI accused it of distilling from o1. In February 2026, [Anthropic alleged](https://fortune.com/2026/02/24/anthropic-china-deepseek-theft-claude-distillation-copyright-national-security/) that DeepSeek, along with other Chinese labs, used fake accounts to harvest millions of Claude conversations for the same purpose. None of this is proven in court, and distillation is common across the industry, but it is a fair part of the picture: part of why DeepSeek was cheap is that it could learn from models others paid to build first. ## Is DeepSeek Open Source? Open Weights vs Open Source DeepSeek is described as open source almost everywhere, including by people who should know better, and the label is only half right. **DeepSeek's models are open weights, not open source.** Since R1, DeepSeek has released its flagship models under the MIT License, one of the most permissive licenses there is, with no strings attached. That is genuinely more open than Kimi, whose [modified license](https://geotoolbox.ai/blog/what-is-kimi-ai) adds an attribution requirement. You can download a DeepSeek model, run it, fine-tune it, and ship it commercially. What you cannot do is rebuild it. Open weights means the finished model files are public. Open source, in the strict sense, would also mean releasing the training data and enough of the recipe to reproduce the model from scratch. DeepSeek publishes detailed papers, more than most, but it does not release its training data, so you can use the model freely without being able to fully audit or recreate how it was made. That distinction is not pedantic, because it changes what you can actually trust and control. The same wave of open-weight models, DeepSeek, Kimi, [Zhipu's GLM](https://geotoolbox.ai/blog/what-is-glm-5-2), [Qwen](https://geotoolbox.ai/blog/what-is-qwen), Llama, Mistral, and now the US-built [Inkling](https://geotoolbox.ai/blog/inkling-ai), is reshaping the market precisely because the weights are free to run, even though none of them are open in the way the word implies. If you are weighing DeepSeek against the rest, our guide to [how the major Chinese models compare](https://geotoolbox.ai/blog/chinese-ai-models-compared) puts them in one table. And it sets up the most practical point about DeepSeek safety: because the weights are public and MIT-licensed, you do not have to use DeepSeek's servers at all, which changes the privacy math entirely. ## Is DeepSeek Safe? Privacy, Bans, and the China Question This is the question DeepSeek gets asked most, and it has more than one honest answer depending on how you use it. Start with content. Like any major model, DeepSeek has guardrails and can still be confidently wrong, so verify what matters. It also follows Chinese content rules: ask the hosted chatbot about Tiananmen Square or Taiwan and it will refuse or echo the official line, and the R1-0528 update was noted for tightening that further. That censorship is heaviest in the hosted app and lighter in the raw open weights you run yourself. The bigger issue for businesses is data. When you use the free app or the API, your prompts go to DeepSeek's servers in China, which puts them under Chinese jurisdiction and the data-access laws that come with it. That concern is not hypothetical hand-waving: in January 2025, security firm [Wiz found a publicly exposed DeepSeek database](https://www.wiz.io/blog/wiz-research-uncovers-exposed-deepseek-database-leak), unauthenticated and open to the internet, leaking over a million log lines including chat history and API keys. DeepSeek secured it after disclosure, but it was a basic lapse on a service handling sensitive prompts. That is why DeepSeek has been [restricted in many places](https://techcrunch.com/2025/02/03/deepseek-the-countries-and-agencies-that-have-banned-the-ai-companys-tech/), and it is worth being precise about what "banned" means. Most actions target the hosted app on official devices, not a blanket consumer ban: the US Navy, Pentagon, NASA, and Congress, plus a growing list of US states (Texas was first, with more than a dozen following), and government bodies in countries including Australia, Taiwan, South Korea, and India have blocked it on their own systems. Italy went further: its privacy regulator ordered DeepSeek blocked over data concerns, and the app was pulled from Italian app stores, while Germany asked Apple and Google to do the same over data-transfer concerns. For most people in most countries DeepSeek is still freely available; the restrictions cluster around governments, sensitive workplaces, and the EU's stricter data rules. The honest mitigation is the one the open weights make possible. If data residency or censorship is a dealbreaker, you do not have to touch DeepSeek's servers: a team can run the open MIT-licensed weights on its own infrastructure, so prompts never leave the building and the hosted app's behavior does not apply. That shifts the security burden onto you to host it properly, but it is the clearest answer to the China question for sensitive work: self-host rather than send. ## DeepSeek vs ChatGPT, Claude, and Kimi The same category error trips up most comparisons, so fix it first: DeepSeek and Kimi are models you can download and run, while ChatGPT and Claude are mostly products that front closed models you rent (OpenAI's separate gpt-oss models aside). With that straight, here is the practical landscape. For how DeepSeek stacks up against the whole open field, see our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking.
ToolWhat it isStrongest atOpen weights?Rough cost
DeepSeek (High-Flyer)Open-weight model + appReasoning, coding, math, very low costYes (MIT)Very low; can self-host free
ChatGPT (OpenAI)Product fronting GPT modelsGeneral use, images, voice, the widest ecosystemNo (flagship); separate gpt-oss models, yesFree tier; paid from $20/mo
Claude (Anthropic)Product fronting Claude modelsWriting, careful reasoning, long documentsNoFree tier; paid from $20/mo
Kimi (Moonshot)Open-weight model + app (also China)Agentic multi-step work, coding, long contextYes (modified MIT)Low; can self-host free
On capability, the honest read is that DeepSeek competes hardest on reasoning, coding, and price, where R1-class models do work comparable to far pricier American ones. Where the closed products lead is general polish, multimodal range (images, voice), reliability across varied tasks, and ecosystem depth. Independent reviewers also flag weaker safety guardrails on DeepSeek, meaning it is easier to push into answering things the closed models refuse. Against [ChatGPT and Claude](https://geotoolbox.ai/blog/claude-vs-chatgpt) the trade is cost and openness versus polish and support; against [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai), DeepSeek is the reasoning-and-cost specialist while Kimi leans more agentic, and both carry the same China-hosting caveats. The distillation questions covered earlier apply here too: part of DeepSeek's value is that it caught up fast and cheap, with help, intended or not, from the models it competes with. ## Is DeepSeek Free? Pricing and How to Access It Yes, DeepSeek is free to use through the web app and mobile apps, no payment required. Where it really stands out is the API, which is priced well below the American labs, even after an August 2026 price increase. The figures below are the [current V4 API rates](https://api-docs.deepseek.com/quick_start/pricing); they change, and DeepSeek retired its older `deepseek-chat` and `deepseek-reasoner` aliases on July 24, 2026, 15:59 UTC, so call `deepseek-v4-flash` and `deepseek-v4-pro` directly.
How you use itPrice (as of August 2026)What you get
Free chat$0Web and mobile app access, no setup
API (V4-Flash)$0.22 in / $0.66 out off-peak, up to $0.44 / $1.32 at peak, per 1M tokensThe cheap, fast workhorse; 1M context
API (V4-Pro)$0.66 in / $1.98 out off-peak, up to $1.32 / $3.96 at peak, per 1M tokensDeepSeek's highest-capability model for reasoning and agents; 1M context
Self-hostFree license; substantial hardware costRun the open weights yourself; data stays on your servers
The change DeepSeek had been warning about since July landed on August 16, 2026: the flat all-day rate was replaced with the peak/off-peak split shown above, peak running 01:00-04:00 and 06:00-10:00 UTC at exactly double the off-peak rate. DeepSeek's [official pricing page](https://api-docs.deepseek.com/quick_start/pricing) is still the source to check before you build, since these rates have already moved twice this year. For comparison, those API prices are a small fraction of what the leading closed models charge, which is the whole reason developers reach for DeepSeek on cost-sensitive, high-volume work. Access comes in four flavors: the free web and app, the OpenAI-compatible API for building, and self-hosting the open weights if you want full control. Self-hosting is realistic mainly for companies with serious GPU hardware, since the flagship models are large, but it is the route that sidesteps the China-hosting concern entirely. For everyone else, the free app or a third-party provider that hosts DeepSeek outside China is the practical middle ground. ## What DeepSeek Means for Your AI Visibility Step back from the specs and there is a marketing angle hiding in DeepSeek's story, and it is sharper than it looks. Every capable new model is another place a customer might ask "what is the best tool for X" or "is [your company] any good" and act on the answer. DeepSeek is one of those places, and the open-weight twist makes it bigger: because the weights are public, DeepSeek does not only answer inside its own app. It gets hosted, fine-tuned, and embedded into a long tail of downstream products you will never audit one by one. Researching this piece surfaced the cleaner lesson. Ask a current AI model "what is DeepSeek" and we found several still describing it as a V3-and-R1 story, with no idea V4 exists. The models can be stale even on their own competitors. That is the whole game in miniature: AI answers are only as current and accurate as the sources they can find, so a clear, up-to-date page is how you get represented correctly instead of through whatever outdated mush a model happens to hold. That is the [foundation of generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization). In practice it splits into two jobs. First, reachability: every one of these models and the crawlers feeding them has to be able to fetch your site, or you are invisible to the live half of the system. Second, consistency: the brands that get described correctly are the ones whose facts line up across the pages a model is likely to read. In our experience at geotoolbox, the businesses that surface well in AI answers are rarely the ones with the prettiest homepage; they are the ones a model can find, parse, and trust without tripping over contradictions. Our guides on [what GEO is](https://geotoolbox.ai/blog/what-is-geo) and [tracking your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) go deeper on both. The open-model wave does not change the playbook so much as widen the field: there are simply more engines that can mention, or mangle, what you have built. The first move is to check whether they can even read you. Run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether the AI crawlers can reach and parse your site, and where the gaps are, before the next model launches and the question gets asked again. ## Frequently Asked Questions ### Is DeepSeek safe to use? It depends on how you use it. The models have guardrails like any assistant, but DeepSeek follows Chinese content rules and censors politically sensitive topics, and its safety filters are generally weaker than the closed leaders'. On data, the hosted app and API send your prompts to servers in China, and DeepSeek had a real exposed-database incident in early 2025, so for sensitive work the safer route is to self-host the open weights rather than use the app. ### Is DeepSeek free? Yes. The web and mobile apps are free to use, and the API is among the cheapest of any major model, roughly $0.22 per million input tokens off-peak for the V4-Flash model as of August 2026 (up to $0.44 during DeepSeek's peak hours). The open weights are also free to download and run under the MIT License if you have the hardware. Prices change, so check the official pricing page. ### Is DeepSeek banned in the US? Not for ordinary consumers. The US bans are targeted: the Navy, Pentagon, NASA, Congress, and a growing list of US states (Texas was the first) have blocked the hosted app on official devices, and other governments abroad have done the same. Italy went further and removed the app from its app stores. But there is no blanket US ban, and most people can still use it freely. ### Is DeepSeek really open source? Not in the strict sense. DeepSeek releases its models as open weights under the permissive MIT License, so you can download, run, and adapt them freely, which is genuinely open. But it does not release its training data or full recipe, so you cannot fully reproduce or audit how the models were built. "Open weights" is the accurate term. ### Did DeepSeek really cost $6 million to build? No. The $6 million figure (about $5.576 million) was the cost of a single final training run for the V3 model, not the cost of the company. It leaves out research, failed runs, staff, and hardware. Independent analysis estimates DeepSeek's real infrastructure at around 50,000 GPUs and well over a billion dollars. The efficiency was real, but the headline number was the smallest true figure available. ### Is DeepSeek better than ChatGPT? It depends on the task. DeepSeek is praised for reasoning, coding, and very low cost, and competes closely with the top models there. ChatGPT is broader and more polished, with images, voice, and a wider ecosystem, and stronger safety guardrails. Many people use both and pick per job. ## Sources - DeepSeek V4 preview announcement - DeepSeek, April 2026 - `api-docs.deepseek.com/news/news260424` - DeepSeek API pricing - DeepSeek (official) - `api-docs.deepseek.com/quick_start/pricing` - DeepSeek-V4-Pro model card - DeepSeek (Hugging Face) - `huggingface.co/deepseek-ai/DeepSeek-V4-Pro` - DeepSeek company overview - Wikipedia - `en.wikipedia.org/wiki/DeepSeek` - Wiz Research uncovers exposed DeepSeek database - Wiz, January 2025 - `wiz.io/blog/wiz-research-uncovers-exposed-deepseek-database-leak` - DeepSeek may have spent ~$1.6B on buildouts (SemiAnalysis) - Tom's Hardware, 2025 - `tomshardware.com/tech-industry/artificial-intelligence/deepseek-might-not-be-as-disruptive-as-claimed-firm-reportedly-has-50-000-nvidia-gpus-and-spent-usd1-6-billion-on-buildouts` - Anthropic accuses Chinese labs of distillation via Claude - Fortune, February 2026 - `fortune.com/2026/02/24/anthropic-china-deepseek-theft-claude-distillation-copyright-national-security` - The countries and agencies that have banned DeepSeek - TechCrunch, 2025 - `techcrunch.com/2025/02/03/deepseek-the-countries-and-agencies-that-have-banned-the-ai-companys-tech` --- ## What Is Google Gemini? Models, Pricing & Features (2026) > What is Google Gemini? A current guide to Google's AI: the models, pricing, how to use it, privacy, Gemini vs ChatGPT, and why it matters for your brand. - Canonical: https://geotoolbox.ai/blog/what-is-gemini - Published: 2026-06-22 · Updated: 2026-08-05 Google Gemini is the AI that now answers in Google Search, lives in the Gemini app, and quietly turned on inside Gmail, Docs, and your Android phone. If you want the plain version of what Gemini actually is, which models are current, what it costs, and whether you can trust it, this is it, current as of July 2026. First, the name. "Gemini" is also the zodiac sign, a 1960s NASA program, and a crypto exchange, and the search results mix them together. This article is about [Google's AI](https://geotoolbox.ai/blog/how-does-ai-search-work), not the constellation, the spacecraft, or the Winklevoss exchange. And there is a part most "what is Gemini" explainers skip: because Gemini powers the AI answers in Google Search, what it says about your company is increasingly what searchers see. We will cover what Gemini is, the current models, pricing, how to use it, the honest take on privacy and safety, how it compares to ChatGPT and Claude, where it lands among [the best AI search engines](https://geotoolbox.ai/blog/best-ai-search-engines), and what all of it means for whether AI mentions your brand. ## What Is Google Gemini?
![Three cards separating Gemini the model family, the chat app, and the assistant layer.](/blog/what-is-gemini/gemini-model-app-assistant.png)
"Gemini" means three different things — the models, the app, and the layer inside Google's products.
**Google Gemini is a family of multimodal AI models built by Google DeepMind, plus the assistant app and features built on top of them.** It is Google's direct answer to ChatGPT and Claude. "[Multimodal](https://geotoolbox.ai/glossary/multimodal-ai)" means it was built to handle text, images, audio, and video together, rather than text alone, so you can paste a screenshot, share a PDF, or talk to it and get a useful answer back. If the name feels new, the product is not. Gemini is the assistant Google launched in 2023 as **Bard**, then [renamed to Gemini in February 2024](https://en.wikipedia.org/wiki/Google_Gemini) when it moved onto the Gemini model family. Bard no longer exists as a separate thing; everything folded into Gemini. The word "Gemini" does double duty, which is the first source of confusion. It can mean the underlying model (the [large language model](https://geotoolbox.ai/glossary/large-language-model) doing the work), the app you chat with at gemini.google.com, or the assistant baked into Google's other products. When Google says "Gemini," context decides which one. Throughout this guide, we will be specific about whether we mean a model, the app, or a feature. The single most important thing to understand up front is reach. ChatGPT is a place you go; Gemini shows up where you already are, inside Search, Gmail, Docs, and Android, whether you went looking for it or not. Whether that is convenient or intrusive depends entirely on how much of your life already runs on Google, which is the real question this article will help you answer. If your work runs on Microsoft instead, our [Copilot vs Gemini](https://geotoolbox.ai/blog/copilot-vs-gemini) comparison weighs the two side by side, and our [best ChatGPT alternatives](https://geotoolbox.ai/blog/chatgpt-alternatives) roundup covers the wider field sorted by job. ## Who Makes Gemini, and Where the Name Comes From **Gemini is made by Google DeepMind, the AI division Google formed in 2023 by merging its two research labs, DeepMind and Google Brain.** That merger is the answer to "who owns Gemini AI": it is Google's own AI, not a startup Google bought and not a partnership. The same group builds the Gemini models, and Google's product teams put them in the consumer app, Search, Workspace, and the developer tools other companies build on. The name is a small Easter egg with three layers. "Gemini" is Latin for "twins," and it is both a zodiac sign and a northern constellation whose two brightest stars, Castor and Pollux, are named for mythological twins. Google leaned on that "twin" idea two ways: the model was built to pair different abilities, like reading text and images at once, and the project itself was the pairing of the two merged research teams. Google has also said the name nods to NASA's Project Gemini, the two-astronaut spacecraft that bridged the early space program and the Apollo missions. That overlap is exactly why the search results are messy. Type "what is Gemini" and you will get the AI, the star sign, and the space program in the same breath. For the rest of this guide, Gemini means Google's AI. ## The Gemini Models, Explained (As of July 2026) **Gemini is not one model; it is a lineup, and Google ships new versions almost every month, which is why so many guides are out of date.** The honest starting point: most explainers you will find still describe Gemini 1.5 or 2.5 as current. They are not. Here is the shipped lineup as of July 2026, straight from [Google DeepMind's model page](https://deepmind.google/models/gemini/).
ModelBest forStatus (July 2026)
Gemini 3.6 FlashFast, everyday answers, agents, and coding; the default you get most of the timeCurrent default model in the Gemini app (July 2026), successor to 3.5 Flash; Google Search's AI Mode still runs on Gemini 3.5 Flash as its default
Gemini 3.1 ProHarder reasoning, complex and creative tasksThe current "Pro" model on paid plans
Gemini 3.1 Deep ThinkThe hardest problems in science, math, and engineering; it weighs several approaches before answeringAvailable on the top tier
Gemini 3.5 Flash-LiteHigh-volume, cost-sensitive tasks that still need decent intelligenceNew budget tier (July 2026), mostly used by developers; succeeds 3.1 Flash-Lite
Gemini 3.5 ProThe next reasoning flagship, with a larger context windowAnnounced; "coming soon" per Google, not yet generally available
A few things help cut through the version soup. **Flash means fast and cheap, Pro means slower and stronger**, and that split holds across every generation. Within Flash itself there is now a further split between the workhorse and the budget tier, which our [Gemini 3.6 Flash vs 3.5 Flash-Lite](https://geotoolbox.ai/blog/gemini-3-6-flash-vs-3-5-flash-lite) guide breaks down. The version numbers do not sort across the two tracks: Gemini 3.5 Flash is newer than 3.1 Pro, but it is not "better than Pro," it is the fast model in a later release. You rarely choose by hand anyway; the app defaults to Flash and reaches for Pro on harder questions or when you are on a plan that allows it. The other number that matters is the [context window](https://geotoolbox.ai/glossary/context-window), which is how much text the model can hold in mind at once. Today's top Gemini models support up to a roughly one-million-[token](https://geotoolbox.ai/blog/what-are-tokens-in-ai) window, enough for a long book or a large codebase in a single go, though the limit you actually get varies by plan and surface, and Google has said the [upcoming 3.5 Pro](https://geotoolbox.ai/blog/gemini-3-5-pro) pushes the ceiling further. If you want the mechanics of how any of these models turn your prompt into an answer, our explainer on [how ChatGPT works](https://geotoolbox.ai/blog/how-does-chatgpt-work) covers the same transformer machinery Gemini uses. Google reprices and renames Gemini models constantly. The shape stays stable, Flash for speed and Pro for reasoning, but the exact version you are handed depends on your plan, your region, and the week. Check Google's own model page before quoting a specific version. ## What Can You Do With Gemini? **Gemini does what you expect from a modern assistant, and then a layer more because it is wired into Google's products.** At the core, it answers questions, writes and edits text, summarizes long documents, writes and debugs code, reads images and screenshots, and holds a back-and-forth conversation. That part is table stakes across ChatGPT, Claude, and Gemini. The differences are at the edges, and most of them come from Google's ecosystem. Inside **Google Workspace**, Gemini drafts and rewrites in Docs, builds formulas and pivot tables in Sheets, generates slides, and summarizes long threads or writes replies in Gmail. In **Google Search**, Gemini powers the [AI Overviews](https://geotoolbox.ai/glossary/ai-overviews) at the top of results and the conversational AI Mode, so most people use Gemini without ever opening the app. It also runs on Android as the default assistant and inside Chrome. On top of chat, Gemini has a set of named features worth knowing: - **Image and video generation** through Google's media models: the viral Nano Banana image models, Veo for video, and the newer [Gemini Omni](https://geotoolbox.ai/blog/gemini-omni), which Google pitches as creating "anything from any input," starting with video - **Deep Research**, an agent that runs many searches and writes up a cited report, which you can point at your own uploaded files - **Gemini Live**, real-time voice conversation that can also see through your camera or screen - **[Gems](https://geotoolbox.ai/blog/gemini-gems)**, custom versions of Gemini you set up for a recurring task, like a writing editor or a study coach Google is also pushing into [agentic tools](https://geotoolbox.ai/glossary/ai-agent) that act, not just answer: [Gemini Spark](https://blog.google/innovation-and-ai/technology/ai/google-io-2026-all-our-announcements/), pitched as a personal agent that takes actions on your behalf, plus the Gemini CLI for developers and Project Mariner for browser tasks. Those are early and change quickly, so treat them as direction rather than finished products. For everyday use, the durable value is the combination of a capable model and the fact that it already lives inside the Google apps you open every day. ## Gemini Pricing and the Free Tier **Yes, Gemini has a genuinely useful free tier, and for a lot of people it is all they need.** The free plan gives you Gemini 3.6 Flash as the default, a daily allotment of the stronger Pro model for harder questions, image generation, voice mode, and a small number of Deep Research reports per month. The catch is throttling: heavy use of the Pro model hits daily caps, and during busy periods you can get bumped to the lighter model. For casual questions you will rarely notice; for all-day, document-heavy work you will. If you outgrow free, the paid plans look like this as of July 2026. Google reshuffled and repriced these tiers more than once in 2026, so treat the figures as the current shape, not a permanent quote, and check [Google's own plans page](https://one.google.com/about/google-ai-plans/) before paying. For a full tier-by-tier breakdown, including the developer API token costs, see our [Gemini pricing guide](https://geotoolbox.ai/blog/gemini-pricing).
PlanPrice (US, July 2026)What you get
Free$0Gemini 3.6 Flash default, daily Pro access, image and voice, limited Deep Research
Google AI Plus$4.99/moHigher limits and more storage; the cheap step up (cut from $7.99 in June 2026)
Google AI Pro$19.99/moGemini 3.1 Pro access, Deep Research, Workspace integration, 5 TB storage
Google AI Ultrafrom $99.99/moHighest limits, Deep Think, Veo video, large storage; top configuration runs higher
The two numbers people miss: **Google AI Plus dropped to $4.99 a month** in June 2026, per [9to5Google](https://9to5google.com/2026/06/08/google-ai-plus-price-drop/), making it one of the cheapest serious AI subscriptions from a major lab, and **Google AI Ultra now starts around $99.99 a month**, after Google added a cheaper entry tier and cut its premium plan from the earlier $249.99. Developers do not use these consumer plans at all; they pay per token through the Gemini API, which is billed separately. The honest recommendation: start on the free tier, use it hard for a couple of weeks, and only pay once you actually hit a wall. Most people do not. ## Gemini on Your Phone, and How to Turn It Off **If Gemini showed up on your Android phone without you installing it, that is by design: Google is replacing Google Assistant with Gemini as the default assistant.** That is why so many people meet Gemini not by choosing it but by long-pressing the power button and finding a new assistant staring back. It is the single most common complaint about Gemini, and it is a fair one. What the phone app actually does is the same as the web version: answer questions, set things up by voice, read what is on your screen if you allow it, and tie into your Google apps. The friction is that it asks for access to do the useful parts, and the prompts can feel pushy. You are not stuck with it. You can switch your default assistant back in your phone's settings under the digital assistant or default apps options, and you can pause or limit what Gemini can reach. On most phones Gemini is a system-level app you cannot fully uninstall, but you can disable it, revoke its permissions, and stop it from being the assistant that pops up. Uninstalling or disabling the app does not delete your account history; that is managed separately in your Google account, which we cover next. The short version: Gemini being on your phone is Google's push, not a virus or a setting you broke, and you can turn the assistant role off without much trouble even if you cannot make the app disappear entirely. ## Is Gemini Safe? Privacy and Your Data **Gemini is safe to use in the ordinary sense, but "safe" and "private" are not the same thing, and the privacy defaults are worth understanding.** By default, your conversations with the consumer Gemini app can be used to improve Google's AI, and a sample of chats may be read by trained reviewers. Google's own guidance is blunt about the implication: do not enter anything confidential you would not want a reviewer to see. You have controls, but you have to use them. You can turn off **Gemini Apps Activity** in your Google account, which stops future chats from being saved to your history, though Google has said human-reviewed conversations can be kept for up to three years even after you turn it off. The bigger flashpoint is Gmail: in late 2025, US users reported that Gemini's smart features in Gmail, Chat, and Meet were enabled by default, an opt-out rather than an opt-in, and a [class-action lawsuit](https://natlawreview.com/article/silent-switch-new-lawsuit-alleges-google-uses-gemini-ai-secretly-read-gmail-chat) alleges the setting was buried. Google disputes the framing, says it did not change the setting, and says it does not use your Gmail content to train Gemini. The practical takeaway: check your smart-features and Personal Intelligence settings yourself rather than assuming they are off. Paid Workspace and enterprise accounts run on separate terms, and that data is not used to train the consumer models. There is also the accuracy side of safety. Like every [large language model](https://geotoolbox.ai/glossary/large-language-model), Gemini can state wrong things with total confidence, and because it powers AI Overviews, those mistakes sometimes show up at the top of Search. Google has had public stumbles, including [pausing Gemini's ability to generate images of people](https://www.theguardian.com/technology/2024/feb/28/google-chief-ai-tools-photo-diversity-offended-users) in early 2024 after it produced historically inaccurate depictions. The practical rule is the same one that applies to any AI: treat it as a fast, capable starting point, and verify anything that actually matters before you rely on it. ## Gemini vs ChatGPT and Claude **All three are excellent in 2026, and the gap between them is now small enough that fit matters more than raw capability.** The right question is not "which is smartest," it is "which one already lives where I work." Here is the honest high-level split. For the full head-to-head, see our [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt) comparison.
PickWhen it is the better fit
GeminiYou live in Gmail, Docs, and Google Search, or you need its very large context window for big documents
ChatGPTYou want the broadest ecosystem of apps and integrations and strong all-round performance
ClaudeYou prioritize careful long-form writing, nuanced analysis, and coding
Gemini's real edge is not a benchmark score; it is the integration. It can summarize and draft from your Gmail, work inside your documents, and answer inside Search without you opening a separate app. For someone whose work already runs on Google, that convenience usually outweighs a few points on a leaderboard. The flip side: if your writing standard is high or you are deep in a non-Google stack, you may prefer a competitor, and many people simply keep two open and switch by task. We dig into the head-to-heads in [Claude vs Gemini](https://geotoolbox.ai/blog/claude-vs-gemini) and [Grok vs Gemini](https://geotoolbox.ai/blog/grok-vs-gemini), and if you are weighing the field more broadly, the brand explainers for [Claude](https://geotoolbox.ai/blog/what-is-claude-ai) and [Grok](https://geotoolbox.ai/blog/what-is-grok) cover those engines the same way this guide covers Gemini. The practical advice: do not agonize. Pick the one that fits your daily tools, learn it well, and you will get more out of it than from constantly chasing the newest release. ## Why Gemini Matters for Marketers and Brands **Here is the part the other "what is Gemini" guides leave out: Gemini is not just a chatbot people visit, it is the engine behind the AI answers in Google Search.** [AI Overviews](https://geotoolbox.ai/blog/google-ai-overviews-seo) and AI Mode are powered by Gemini, which means the model is already summarizing your category, recommending tools, and describing your company to people who never open the Gemini app. What Gemini "knows" about you is becoming what a huge share of searchers see first. The scale is the point: as of mid-2026, [Google says](https://blog.google/products-and-platforms/products/search/new-controls-website-owners/) AI Overviews reach more than 2.5 billion people a month and AI Mode has passed a billion, all powered by Gemini's models. That changes the job. Classic SEO was about ranking a blue link; [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo) is about being the source the AI trusts enough to cite and summarize. The mechanics overlap but they are not identical, which is why we treat [GEO, AEO, and SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo) as related but distinct disciplines. If you want the practical playbook, our guide on [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search), the engine-specific [Gemini SEO](https://geotoolbox.ai/blog/gemini-seo) walkthrough, and the case for a complete, consistent [brand entity](https://geotoolbox.ai/blog/entity-seo) are the place to start. One nuance trips up a lot of people. Google offers a crawler control called **Google-Extended** that lets you opt your content out of training Gemini and grounding its app answers. Google has said it is not a Search ranking signal, and it does not pull you out of AI Overviews, which run on the regular Search index. If you do want out of the AI answers, Google has been rolling out a Search Console control that can pull your site from AI Overviews and AI Mode while keeping your normal Search listing, but opting out means you forfeit any traffic from those AI features, so weigh it carefully. Google-Extended itself protects your content from model training without making you invisible in Search, a trade-off worth understanding before you touch your robots file. The hard part is that you cannot fix what you cannot see. Gemini does not show its sources the way a search results page does, so most brands have no idea whether the AI describes them accurately, recommends a competitor, or invents a detail. That is the gap we built geotoolbox to close: it monitors how Gemini, AI Overviews, ChatGPT, and the other engines represent your brand, which sources they cite, and where you are missing from the answers, so you can act on real data instead of guesses. If you want to see where AI engines mention your competitors but not you, [geotoolbox's Citation Interceptor](https://geotoolbox.ai/features/citation-interceptor) maps exactly that across Gemini, Google AI Overviews, ChatGPT, Perplexity, Claude, and more, so you know which conversations to get into. You can also learn the broader approach in our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). ## Frequently Asked Questions ### How much does Gemini cost per month? Gemini has a free tier, and the paid consumer plans as of July 2026 are Google AI Plus at $4.99 a month, Google AI Pro at $19.99 a month, and Google AI Ultra starting around $99.99 a month, with higher configurations costing more. Google reprices these often, so confirm the current numbers on Google's plans page before subscribing. Developers pay separately, per token, through the Gemini API. ### Is Google Gemini free? Yes. The free tier gives you Gemini 3.6 Flash as the default model, a daily allowance of the stronger Pro model, image generation, voice mode, and a limited number of Deep Research reports. Heavy users hit daily caps and get throttled at busy times, which is the main reason to consider a paid plan. For most people the free tier is enough. ### Why is Gemini on my phone, and how do I turn it off? Google is replacing Google Assistant with Gemini as the default assistant on Android, which is why it can appear without you installing it. You can switch your default assistant back and revoke Gemini's permissions in your phone's settings. On most phones you cannot fully uninstall it because it is a system app, but you can disable it and stop it from being the assistant. ### Is Google Gemini safe to use? It is safe for everyday use, but by default your consumer chats can be used to improve Google's AI and a sample may be reviewed by humans, so do not paste confidential information. You can turn off Gemini Apps Activity to stop saving chats, but note Google switched Gemini's Gmail smart features on by default for US users in late 2025 (an opt-out that drew a class action), so check those settings yourself. As with any AI, verify important facts, because Gemini can be confidently wrong. ### Who owns Gemini? Gemini is made by Google, specifically Google DeepMind, the AI division formed in 2023 by merging DeepMind and Google Brain. It is Google's own AI, and it was previously called Bard before the February 2024 rebrand. ### Is Gemini better than ChatGPT? Neither is universally better in 2026; it depends on fit. Gemini wins if you live in Google's apps or need its very large context window, while ChatGPT has a broader integration ecosystem and Claude is often preferred for long-form writing. The honest move is to pick the one that fits your daily workflow rather than chasing benchmarks. ## Sources - Gemini models - Google DeepMind - `deepmind.google/models/gemini` - Google Gemini - Wikipedia - `en.wikipedia.org/wiki/Google_Gemini` - What is Google Gemini? - IBM - `ibm.com/think/topics/google-gemini` - Google AI plans - Google One - `one.google.com/about/google-ai-plans` - New controls and insights for website owners (AI Overviews reach) - blog.google - `blog.google/products-and-platforms/products/search/new-controls-website-owners` - 100 things we announced at Google I/O 2026 (Gemini Spark, Omni) - blog.google - `blog.google/innovation-and-ai/technology/ai/google-io-2026-all-our-announcements` - Lawsuit alleges Google's Gemini reads Gmail by default - National Law Review - `natlawreview.com/article/silent-switch-new-lawsuit-alleges-google-uses-gemini-ai-secretly-read-gmail-chat` - Google AI Plus gets price drop to $4.99 (June 2026) - 9to5Google - `9to5google.com/2026/06/08/google-ai-plus-price-drop` - Google chief admits 'biased' AI tool's photo diversity offended users - The Guardian - `theguardian.com/technology/2024/feb/28/google-chief-ai-tools-photo-diversity-offended-users` --- ## What Is GLM 5.2? China's Open-Weight AI, and Your Visibility > What is GLM 5.2? Zhipu's open-weight AI explained: the specs, honest benchmarks, the China question, pricing, and what it means for your AI visibility. - Canonical: https://geotoolbox.ai/blog/what-is-glm-5-2 - Published: 2026-06-22 · Updated: 2026-08-18 China's Zhipu AI shipped a frontier-grade open-weight model just days after the US forced Anthropic to take its most powerful Claude models offline, and nearly all the coverage has been about coding benchmarks and price. Here is the plain version of what GLM 5.2 actually is, who builds it, whether it is safe, and where it stands, current as of mid-2026. And then the part the developer write-ups skip: why a model most marketers will never touch directly is still worth a few minutes of your attention, and what the open-weight wave it belongs to means for your visibility in AI answers. On August 14, 2026, Zhipu released **GLM-5.3** on the same GLM Coding Plan tiers (Lite/Pro/Max) covered below, built on the same base as 5.2 through extended post-training rather than a new pretraining run. Zhipu claims roughly a 50% jump in coding capability over 5.2 on its internal evals and a first-place finish among open-weight models on Terminal-Bench 3.0 and Agents' Last Exam. Open weights were not available at launch; Zhipu said they would follow about two weeks later. Everything below about GLM 5.2's specs, pricing, and safety profile still applies to 5.2 itself; treat GLM-5.3 as the newer sibling until weights land and the picture is independently verified. ## What Is GLM 5.2? **GLM 5.2 is the latest flagship open-weight model from Zhipu AI, the Chinese lab that operates internationally as Z.ai.** It launched on June 13, 2026 as a coding-first frontier model with a one-million-token context window and an MIT license, with the downloadable weights and a standalone API following within days. That license means anyone with the hardware can run it. GLM stands for General Language Model, the name Zhipu has used for its [large language models](https://geotoolbox.ai/glossary/large-language-model) since the series began. Like other open models, GLM 5.2 is really three things, and keeping them straight clears up most of the confusion. There are the open weights, which Zhipu publishes so anyone can download, run, and adapt the model. There is the hosted Z.ai chatbot and API, which Zhipu runs on its own servers. And there is the wider ecosystem of third-party tools that have wired GLM 5.2 in. The model is open; the service around it is Zhipu's. That split matters because the answer to most questions about GLM 5.2 is different depending on which version you mean. Where it earned its reputation is doing frontier-grade coding and agent work at a fraction of what the closed American models cost, which is why its arrival made the industry pay attention to a name most marketers had never heard. If you search the bare term "GLM," most results are about the generalized linear model, a staple of statistics. That is a different thing entirely. The AI model is always written with its version number, "GLM 5.2" or "GLM-5.2," and that is what this article is about. ## Who Makes GLM 5.2? Zhipu AI and Z.ai **GLM 5.2 comes from Zhipu AI, a Beijing-based lab founded in 2019 as a spinout from Tsinghua University's Knowledge Engineering Group.** Internationally it goes by Z.ai, and it has become one of China's most aggressive publishers of open-weight models, shipping new flagships every couple of months. That cadence is the backdrop for GLM 5.2: it is the third release in the GLM-5 line in four months. The market noticed. When Zhipu announced GLM 5.2 would be released with open weights, its shares jumped. The [South China Morning Post reported](https://www.scmp.com/tech/tech-trends/article/3357115/zhipu-ais-stock-rockets-after-chinese-firm-makes-glm-52-open-source) that the stock surged as much as 48 percent intraday and closed up 32.8 percent on June 15, 2026. Investors read a giveaway as a strategic win, not a cost, which tells you how the open-weight race is being valued right now. It helps to place Zhipu against the other Chinese labs you have probably seen in headlines, because they are easy to blur together. Zhipu and Z.ai are the GLM company, out of Beijing and rooted in Tsinghua. That is a different outfit from [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek), the Hangzhou lab funded by a quant hedge fund, and from Moonshot, the company behind [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai). All three publish open weights, all three operate under Chinese law, and all three are part of the same wave putting cheap, capable models within everyone's reach. If you want them lined up side by side, our [comparison of the major Chinese AI models](https://geotoolbox.ai/blog/chinese-ai-models-compared) does exactly that. ## The Specs: What Actually Changed in 5.2
![Bar chart showing GLM 5.2's context window jump from roughly 200K tokens to one million.](/blog/what-is-glm-5-2/glm-context-window-jump.png)
The real upgrade from GLM 5.1 to 5.2: a five-fold jump in how much text the model can hold at once.
**The headline change in GLM 5.2 is not a bigger brain, it is a bigger window.** Zhipu pushed the [context window](https://geotoolbox.ai/glossary/context-window) from roughly 200,000 tokens in GLM 5.1 to a usable one million, and it [describes that capacity](https://z.ai/blog/glm-5.2) as engineering-usable rather than a number on a spec sheet. In plain terms, you can hand the model an entire mid-sized codebase or a long stack of documents at once, around 750,000 words, and it can reason over the whole thing without chunking. Under the hood, GLM 5.2 is a Mixture-of-Experts model with around 753 billion total parameters, of which only about 40 billion are active for any given [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), per its [Hugging Face model card](https://huggingface.co/zai-org/GLM-5.2). (Some write-ups cite roughly 744 billion; the count from the published weights is a little higher.) That total is barely changed from GLM 5.1, so the jump is in the context window, not the model's size. The Mixture-of-Experts design is the trick behind the low running cost: it activates only a slice of those parameters per token, so each token takes far less compute than the headline size suggests, even though the full model is still heavy to host. The release also adds two thinking-effort levels, High and Max, so you can trade speed for deeper reasoning when a task needs it. One correction worth making, because a few early write-ups got it wrong: GLM 5.2 is a text model. It reads and writes text and code, and the launch materials and model card cover only that. Treat it as text-only until Zhipu publishes anything saying otherwise. Here is how it compares to the model it replaces.
Spec (as of July 2026)GLM 5.2GLM 5.1
ReleasedJune 13, 2026April 7, 2026
Total parameters~753 billion~754 billion
Active per token~40 billion~40 billion
Context window1,000,000 tokens~200,000 tokens
Max output131,072 tokens~128,000 tokens
Reasoning modesHigh and MaxSingle mode
LicenseMIT (open weights)MIT (open weights)
ModalityText onlyText only
## Is GLM 5.2 Open Source? Open Weights vs Open Source **GLM 5.2 is open weights, not open source, and the difference is worth getting right.** Zhipu released the model under the [MIT License](https://huggingface.co/zai-org/GLM-5.2), one of the most permissive there is. You can download the weights, run them on your own hardware, fine-tune them on your own data, and ship the result commercially, with essentially no strings attached. That is genuinely open, and more permissive than some rivals whose licenses add conditions. What you cannot do is rebuild it. Open weights means the finished model files are public. Open source, in the strict sense, would also mean publishing the training data and enough of the recipe to reproduce the model from scratch. Zhipu does not release its training data, so you can use the model freely without being able to fully audit or recreate how it was made. "Open weights" is the accurate term, even though plenty of coverage calls it open source. That distinction is not pedantry, because it changes where GLM 5.2 ends up. Since the weights are public and the license is permissive, the model does not only answer questions inside Zhipu's own app. It gets downloaded, fine-tuned, rebranded, and quietly built into other products you will never see. Hold onto that point. It is the part that actually matters for your brand, and we come back to it at the end. ## How Good Is It, Really? The Honest Benchmark Picture **By the benchmarks that have landed, GLM 5.2 is one of the strongest open-weight models available, and it is still not the strongest model overall.** Both things are true, and the gap between them is the whole story. The picture shifted in late July 2026: Kimi K3's weights went public on July 27, and its 57 on Artificial Analysis's Intelligence Index now leads GLM 5.2's 51 among open models. But at 2.8 trillion parameters K3 is out of self-hosting reach, so GLM 5.2 keeps the #1 spot in our [best open-source LLMs](https://geotoolbox.ai/blog/best-open-source-llms) ranking as the strongest open model you can actually download and run. It is worth being skeptical of the launch hype, because Zhipu actually shipped GLM 5.2 with no published benchmarks at all. As [MarkTechPost noted](https://www.marktechpost.com/2026/06/14/z-ai-launches-glm-5-2-with-a-usable-1m-token-context-two-thinking-effort-levels-and-no-benchmarks-at-launch/), there were no SWE-bench, Terminal-Bench, or Code Arena numbers at release, which left the early "beats GPT-5" claims as vendor assertions until the technical card and independent tests filled the gap days later. When the numbers landed, they told a consistent story. On its headline coding benchmarks, like SWE-bench Pro, GLM 5.2 edges out GPT-5.5 and lands within a few points of [Claude Opus 4.8](https://geotoolbox.ai/blog/what-is-claude-ai), though the two trade places on other coding tests. On harder, more abstract reasoning, it trails both. Here is the picture using Zhipu's own reported scores against the closed-source leaders as they stood at the time of testing. These are vendor-reported numbers rather than an independent cross-lab run, and benchmark setups differ, so read them as directional.
Benchmark (as of July 2026)GLM 5.2Claude Opus 4.8GPT-5.5Gemini 3.1 Pro
SWE-bench Pro (coding)62.169.258.654.2
Terminal-Bench 2.1 (agentic)81.085.084.074.0
FrontierSWE (long-horizon)74.475.172.639.6
HLE (hard reasoning)40.549.841.445.0
The pattern is clear: a coding and agents specialist that leads the open field on coding but sits behind the closed frontier on raw reasoning. Independent platform [Artificial Analysis](https://artificialanalysis.ai/models/glm-5-2) scores it 51 on its Intelligence Index — the top open-weight model it tracked until Kimi K3 arrived at 57 in late July 2026 — but flags a real catch. GLM 5.2 is token-hungry: it generated 140 million tokens to complete the same evaluation that the average comparable model finished in 110 million, so the cheap sticker price buys a model that talks more to get there. Early testers piled on a second caveat, noting it can stumble on deceptively simple tasks and sometimes emulates reasoning rather than achieving it. Strong, cheap, and genuinely useful, then, but not magic. ## Is GLM 5.2 Safe? The China and Data Question **This is the question GLM 5.2 gets asked most, and the honest answer turns on one distinction: who makes the model versus who hosts your data.** A model's national origin matters for your privacy only when the maker is also the company running your prompts. Get that straight and most of the worry sorts itself out. When you use the hosted Z.ai chatbot or API, your prompts travel to Zhipu's servers, which puts them under Chinese jurisdiction and the data-access laws that come with it. For sensitive code or proprietary documents, that is a genuine consideration, and it is the reason some companies will not route that work through the hosted service at all. The same caution applies to any Chinese-hosted model, which is why we walk through it the same way for [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek). Here is the part that changes the math. Because the weights are public and MIT-licensed, you do not have to use Zhipu's servers. A team can download GLM 5.2 and run it on its own infrastructure, so prompts never leave the building at all. Or it can use a non-Chinese provider that hosts the open weights, which keeps the traffic off Zhipu's servers, though your data then sits with that provider instead. Self-hosting shifts the security burden onto you, but it is the clearest answer to the China question for sensitive work: self-host rather than send. There is also a content angle. Like other Chinese-hosted assistants, the official app follows local content rules on politically sensitive topics, and that filtering is heaviest in the hosted product and lighter in the raw weights you run yourself. ## How Much GLM 5.2 Costs and How to Access It **Cost is GLM 5.2's sharpest edge.** The fastest way in is the GLM Coding Plan, a subscription that runs about $10 a month for the Lite tier, $30 for Pro, and $80 for Max, billed quarterly. For developers building on it directly, the standalone [API is priced](https://openrouter.ai/z-ai/glm-5.2) around $1 per million input tokens and roughly $4 per million output, a small fraction of what the leading closed models charge. Depending on which comparison you draw, that lands GLM 5.2 at somewhere between a sixth and a tenth of the cost of comparable frontier access, against the roughly $200 a month a top Claude plan runs. The one asterisk is the token appetite from the benchmark section. Because GLM 5.2 generates roughly a quarter more tokens to finish a job, heavy agentic use eats into that headline saving, so the real cost depends on your workload, not just the rate card. It narrows the gap rather than erasing it. Here are the main ways to reach it.
How you use itCost (as of July 2026)What you get
Z.ai chatbotFreeWeb chat, no setup, quickest way to try it
GLM Coding Plan (Lite)~$10/monthUse it inside a coding tool, roughly 400 prompts/week
GLM Coding Plan (Max)~$80/monthHeavy agentic use, roughly 8,000 prompts/week
Standalone API~$1 in / ~$4 out per 1M tokensBuild GLM 5.2 into your own product
Self-host (open weights)Free license, your own GPUsData stays on your servers; the model is over 1.5 TB, so plan for serious hardware
The open weights live on [Hugging Face](https://huggingface.co/zai-org/GLM-5.2) and ModelScope, and because Zhipu built it to work with the major AI tooling, developers can drop GLM 5.2 into the coding tools they already use with close to a one-line change. That 1.5 TB is also starting to feel less absolute. In July 2026 an open-source engine called [Colibri](https://www.tomshardware.com/tech-industry/artificial-intelligence/colibri-proof-of-concept-gains-frontier-level-1-5-tb-ai-model-novel-approach-runs-on-only-25gb-of-ram-and-shows-promise-for-local-ai-setups) showed GLM 5.2 running on a consumer machine with about 25 GB of RAM, by keeping the model's dense layers in memory and streaming its thousands of routed experts from disk on demand. It is a proof-of-concept, not a usable tool, glacially slow at around a token per second at best, but it points at the direction of travel: the hardware bar for a frontier open-weight model is dropping, which only widens the long tail of tools and services that can quietly run one, and that widening is the part that matters for your brand. ## What Actually Bites When You Deploy It A month of real use has surfaced friction that the launch benchmarks do not capture. None of it is fatal, but all of it costs somebody a day. **The 1M-context speed story does not survive llama.cpp.** GLM 5.2's long-context efficiency comes from sparse attention with a shared indexer layer. Stock llama.cpp originally refused to load the converted weights at all because it expected an indexer tensor on every layer; the fix made those tensors optional, with the commit itself noting the indexer is not yet supported. So llama.cpp loads the model and produces coherent output, but falls back to full attention. You get the model without the mechanism that makes its context window affordable, and long-context prefill costs far more than the spec sheet implies. The feature request has been open since four days after launch. **Tool calling breaks across serving stacks.** On vLLM, setting `tool_choice: required` makes the tool call come back as raw text in `content` instead of `tool_calls`; `auto` works. In Cursor, GLM 5.2 tool calls have been printing as raw XML and terminating the conversation, which Cursor staff confirmed as a GLM-5.2 parsing problem and warned may take some time to fix. This is a repeat of an identical GLM-5.1 issue that Z.ai patched with a chat-template update, so anyone reusing a 5.1-era template or parser config inherits the old failure. The workaround people report is `--chat-template-content-format=string`. **The shipped chat template defaults reasoning effort to max.** The template resolves anything that is not explicitly `high` to `max`, which means an unset or misspelled value silently picks the most expensive setting, the opposite of the OpenAI and Anthropic convention. That matters because High reportedly delivers roughly 95% of Max's quality at about half the output tokens. The default costs you roughly double for a few percent. That last point connects to a wider complaint: GLM 5.2 is token-hungry. Users report burning through subscription quota far faster than expected, in one case at around five times the rate of a comparable Claude plan. Treat the direction as well-attested and the magnitude as anecdotal, since other users push back and attribute it to workflow rather than the model. **The Coding Plan has multipliers the price tag does not mention.** Quota burns at 3x during peak hours (14:00 to 18:00 UTC+8) and 2x off-peak, with a promotion dropping off-peak to 1x through the end of September 2026, so the cheap era has a published expiry. Quotas are weekly on a rolling cycle from your purchase date, not calendar-monthly, and the MCP tool quota is a separate and much smaller allowance that people routinely confuse with the prompt quota. Subscriptions are non-refundable once purchased. And the error you will most likely hit, `1113 Insufficient Balance`, is a catch-all thrown for calling a non-whitelisted model or the wrong endpoint, not necessarily an actual balance problem. ## GLM 5.2 vs Claude, GPT-5.5, DeepSeek and Kimi So which should you actually use? For most teams the answer is not one model but a split: a frontier model like [Claude or GPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) for the hardest tenth of the work, and GLM 5.2 for the cheap, high-volume rest. It earns that role honestly, leading the open field on coding and beating GPT-5.5 on benchmarks like SWE-bench Pro, while the closed leaders keep their edge on the hardest reasoning.
ModelWhat it isStrongest atOpen weights?Rough cost
GLM 5.2 (Zhipu / Z.ai)Open-weight model + appCoding, agents, huge context, low costYes (MIT)Very low; self-host free
Claude Opus 5 (Anthropic)Product fronting a closed modelHardest reasoning, long documents, writingNoPremium
GPT-5.5 (OpenAI)Product fronting a closed modelBroad general use, reasoning, ecosystemNoPremium
DeepSeek V4 (High-Flyer)Open-weight model + appReasoning, math, very low costYes (MIT)Very low; self-host free
Kimi (Moonshot)Open-weight model + appAgentic work, long contextYes (Modified MIT on K2; custom on K3)Low on the K2 line; K3 is priced at the frontier
The timing of the release said as much as the specs. GLM 5.2 went public a day after Anthropic pulled its [Fable 5 and Mythos 5 models](https://geotoolbox.ai/blog/fable-5-ban) offline (a suspension later lifted, with Fable 5 restored globally on July 1, 2026) to comply with a US export-control order aimed at foreign nationals. Zhipu framed its launch as a direct counter: a frontier-class model under an MIT license with no regional limits, pitched as insurance against any one nation or vendor controlling foundational AI. Whatever you make of the politics, it is part of a broader wave of capable Chinese open releases, alongside [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek), [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai), and [Qwen](https://geotoolbox.ai/blog/what-is-qwen), and that wave is the thing with real consequences for your brand. ## What GLM 5.2 Means for Your Brand's AI Visibility Here is the question most marketers actually have when a model like this lands: do I need to do something about it? The short answer is no, not directly, and the reason is a distinction worth holding onto. **GLM 5.2 is a model, not a destination.** You do not "optimize for GLM 5.2" the way you optimize a page for Google. Z.ai does run a chatbot, but for most teams it is not where customers research vendors, the way they increasingly use ChatGPT, Perplexity, Gemini, and Claude. Those products are what you track, and a single model launching changes none of that playbook on its own. It is worth saying plainly, because the myth runs deep: ranking number one on Google does not mean an AI engine will mention you. That is a separate system with its own logic, and a new model does not move it either way. So why bring it up at all? Because the open-weight wave GLM 5.2 belongs to does change one thing: the number of places an answer about you can appear. Two effects are real. First, because the weights are public and the price is near zero, models like this get downloaded, fine-tuned, and built into a long tail of downstream tools, chatbots, and features you will never audit one by one, and each is another surface where your brand can be mentioned or mangled. GLM 5.2 itself is a weak example of that, since a model wired into a developer's editor rarely gets asked which CRM is best, so the concern is really about the wave, not this one coding-focused model. Second, like any model, GLM 5.2 has a fixed knowledge cutoff and the raw weights do not browse the live web, so whether an engine built on it can even see your newest page depends on the product wrapped around it. And because its training data is undisclosed, it may draw on a different mix of sources and regional coverage than a US-built model, which can shade how it describes your market. We watched this play out firsthand. Nine days after GLM 5.2 launched, as an informal check, we asked two AI models with web search off what it was. The newest, GPT-5.5, was honest about its limits and said it had no reliable information, noting the model "may be a model released after sources I can verify." An older model, GPT-4o, confidently invented an answer: it told us GLM 5.2 was built by Alibaba's DAMO Academy, released in September 2023, with 10 trillion parameters and multimodal features. Every one of those details is wrong. When we then checked which brands ChatGPT pulls in around the term, it cross-wired the Z.ai model with an "Alibaba flagship model" and an "Anthropic frontier model." There are two failure modes there, and your brand can hit either. A model that has never heard of you leaves you out, or, like GPT-4o here, makes you up wholesale. A model that half-knows you, recognizing the name but missing clean facts, is the more insidious case: it fills the gaps with confident guesses. That second mode is the same mechanism behind [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations) about your company, and it is why [how an AI engine chooses its sources](https://geotoolbox.ai/blog/how-does-ai-search-work) matters more to you than your Google rank does. You cannot touch the weights, and you cannot tune for any single model. What you can do is make sure every engine and crawler can actually reach your site, and that your facts line up across the pages a model is likely to read. That combination, reachability plus consistency, is the foundation of [generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization). In our experience at geotoolbox, the businesses that surface well in AI answers are rarely the ones with the prettiest homepage; they are the ones a model can find, parse, and trust without tripping over a contradiction. Our guides on [what GEO is](https://geotoolbox.ai/blog/what-is-geo) and [tracking your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) go deeper, and since being cited by one engine never guarantees the next, that work pays off across all of them. GLM 5.2 is worth understanding, but the open-model wave it belongs to does not change your job so much as widen the field, more places an answer about you can show up, and more places it can be wrong. The first move is also the cheapest one, and it is the same no matter which model launches next: find out whether the AI crawlers can even reach and read your site, because everything else depends on it. You can run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) from geotoolbox to see where those gaps are before the next model launches and the question comes around again. ## Frequently Asked Questions ### Is GLM 5.2 the same as a "generalized linear model"? No. GLM the AI model stands for General Language Model and comes from the Chinese lab Zhipu AI (Z.ai). The generalized linear model is an unrelated statistics method that shares the initials. The version number is the tell: when people say "GLM 5.2," they mean the AI model. ### Who makes GLM 5.2, and is it Chinese? Zhipu AI, a Beijing-based lab that operates internationally as Z.ai and was founded in 2019 as a spinout from Tsinghua University. Yes, it is a Chinese company, and it operates under Chinese law. That mainly matters for the hosted API, which is run by a Chinese company and puts your prompts under Chinese jurisdiction. ### Is GLM 5.2 safe to use for business? It depends on how you use it. The hosted Z.ai API routes your prompts to servers under Chinese jurisdiction, which is a real consideration for sensitive or proprietary work. Because the weights are MIT-licensed, a safer route for that work is to self-host so data never leaves your infrastructure. As with any model, it has guardrails and can still be confidently wrong, so verify anything that matters. ### Is GLM 5.2 free, and how much does it cost? The Z.ai chatbot is free to use. The GLM Coding Plan starts at about $10 a month for Lite and runs to about $80 for Max, and the standalone API lists at $1.40 per million input tokens and $4.40 per million output on Z.ai as of July 2026, though third-party providers often run cheaper. Note that Coding Plan quota burns at 2x off-peak and 3x during peak hours, with a promotional 1x off-peak rate running through the end of September 2026. The open weights are free to download and run under the MIT License if you have the hardware, since the full model is over 1.5 TB. ### Is GLM 5.2 better than ChatGPT or Claude? On coding, it is competitive: it beats GPT-5.5 on SWE-bench Pro and lands within a few points of Claude Opus 4.8 on agentic tasks. On the hardest abstract reasoning it trails both. It is the strongest open-weight model you can realistically self-host, though not the overall best; since late July 2026 Kimi K3 has passed it on raw benchmark scores. Many teams therefore use GLM 5.2 for high-volume work and keep a closed model for the toughest problems. ### Does GLM 5.2 power an AI search engine I need to show up in? Not directly. GLM 5.2 is a coding-and-agents model with a chatbot, not a consumer AI-search engine like Perplexity, so there is nothing to "optimize for." You track the products your buyers actually use, ChatGPT, Perplexity, Gemini, and Claude, and you keep your brand reachable and consistent across the web so any engine, including ones built on open models like this, can describe you correctly. ## Sources - GLM-5.2: Built for Long-Horizon Tasks - Z.ai (official announcement), June 2026 - `z.ai/blog/glm-5.2` - GLM-5.2 model card - Zhipu AI (Hugging Face) - `huggingface.co/zai-org/GLM-5.2` - Z.ai GLM 5.2 API pricing - OpenRouter - `openrouter.ai/z-ai/glm-5.2` - GLM-5.2 model analysis - Artificial Analysis - `artificialanalysis.ai/models/glm-5-2` - Zhipu AI's stock rockets after it makes GLM-5.2 open source - South China Morning Post, June 2026 - `scmp.com/tech/tech-trends/article/3357115/zhipu-ais-stock-rockets-after-chinese-firm-makes-glm-52-open-source` - Z.ai launches GLM-5.2 with a usable 1M-token context and no benchmarks at launch - MarkTechPost, June 2026 - `marktechpost.com/2026/06/14/z-ai-launches-glm-5-2-with-a-usable-1m-token-context-two-thinking-effort-levels-and-no-benchmarks-at-launch` - Colibri proof-of-concept runs GLM-5.2's 1.5-TB model on 25GB of RAM - Tom's Hardware, July 2026 - `tomshardware.com/tech-industry/artificial-intelligence/colibri-proof-of-concept-gains-frontier-level-1-5-tb-ai-model-novel-approach-runs-on-only-25gb-of-ram-and-shows-promise-for-local-ai-setups` --- ## What Is Grok? Elon Musk's xAI Chatbot, Explained > What is Grok? A current guide to xAI's chatbot: who makes it, the models, pricing, Grok vs ChatGPT, the controversies, and what it means for your brand. - Canonical: https://geotoolbox.ai/blog/what-is-grok - Published: 2026-06-22 · Updated: 2026-08-14 Grok is the AI chatbot that argues with people on X, generates images and video, and periodically makes the news for saying something it should not have. If you want the plain version of what Grok actually is, who builds it, what it costs, and whether you can trust it, this is it, current as of August 2026. Most "what is Grok" articles stop at "it is Elon Musk's ChatGPT rival." That is true, but it misses the part that matters if you run a website or a brand: Grok answers questions using live posts on X and the open web, which means it has its own opinion about your company, and that opinion is increasingly what people see. We will cover what Grok is, the model lineup, how to access it, how it compares to ChatGPT, the honest picture on accuracy and safety, and what all of it means for whether AI mentions you.
![Three-step flow showing Grok answering questions from live X posts and web search.](/blog/what-is-grok/grok-live-retrieval-flow.png)
Grok's freshness comes from live retrieval over X and the web, not from the model's memory.
## What Is Grok? **Grok is a conversational AI assistant built by xAI, the company Elon Musk founded in 2023 (now operating as SpaceXAI after the 2026 SpaceX merger).** It answers questions, writes and debugs code, analyzes images, and generates its own images and video, in the same general category as ChatGPT, [Google Gemini](https://geotoolbox.ai/blog/what-is-gemini), and Claude. Its defining trait is live search: in its main surfaces, the X app and grok.com, Grok pulls from public posts on X (formerly Twitter) and the open web, so it can answer about something that happened minutes ago. The model underneath still has a training cutoff; that freshness comes from its search tools, not from the model's memory. Two things set Grok apart from the rest of the field. The first is that real-time connection to X, which makes it useful for breaking news, live sentiment, and "what are people saying right now" questions that stump models without live retrieval. The second is tone. xAI markets Grok as "truth-seeking" and deliberately built it to be less filtered and more willing to answer provocative questions than its competitors, with a sarcastic, slightly rebellious personality. That positioning is the source of both its appeal and most of its controversies. Grok is also a [large language model](https://geotoolbox.ai/glossary/large-language-model) product in the fullest sense: it runs as a standalone app and website, lives inside X, ships in Tesla vehicles, and exposes an API for developers. It is a family of models reached through several front doors, not a single thing you use one way. ## Who Makes Grok, and Why It Is Called "Grok" **Grok is made by xAI, the artificial intelligence company Elon Musk founded in 2023.** Musk started xAI after parting ways with OpenAI, which he had co-founded, and pitched Grok as a less restricted alternative to assistants he argued had become too cautious. The first version launched in November 2023 as a perk for X subscribers. The corporate structure has shifted twice since: [xAI acquired X](https://en.wikipedia.org/wiki/XAI_%28company%29) (the social platform) in 2025, and in February 2026 SpaceX acquired xAI, so Grok now sits inside Musk's wider SpaceX group rather than as a standalone company. The name comes from science fiction, not from an acronym. "Grok" was coined by Robert A. Heinlein in his 1961 novel [Stranger in a Strange Land](https://en.wikipedia.org/wiki/Stranger_in_a_Strange_Land), where it is a Martian verb meaning to understand something so completely that you merge with it. It entered general English to mean deep, intuitive understanding, which is the meaning Musk was reaching for. Three different things share the name, which is where the confusion starts. There is **Grok**, the xAI chatbot this article is about. There is **Groq**, spelled with a q, a completely separate company, founded in 2016, that makes AI inference chips; it publicly objected when xAI launched with the near-identical spelling. And there is **grok** the lowercase developer term, a pattern-matching syntax used to parse log files in tools like Logstash. If you searched for one and landed on another, that is why. For the rest of this article, Grok means xAI's chatbot. ## The Grok Model Lineup, From Grok 1 to Grok 4.6 **Grok is not one model but a fast-moving family, and xAI ships new versions faster than most explainers can keep up.** The pattern has been a major release roughly every few months, each adding reasoning, longer context, or new modes. Here is [the lineup that matters](https://en.wikipedia.org/wiki/Grok_%28chatbot%29), with the dates, as of August 2026.
ModelReleasedWhat changed
Grok-1November 2023The first version, launched inside X; weights later open-sourced under Apache 2.0 in March 2024
Grok-1.5March 2024Longer context and a vision variant (Grok-1.5V, April 2024) that could read images
Grok-2August 2024Stronger general model; added image generation; later released as open weights in 2025
Grok-3February 2025First reasoning model, with a "Think" mode and DeepSearch for multi-step web research
Grok-4July 2025Reasoning-first and multimodal; introduced the "Grok 4 Heavy" tier that runs several agents in parallel on hard problems
Grok 4.1 and later point releasesNov 2025 - 2026Grok 4.1 (and a faster "4.1 Fast") shipped November 2025, with further point releases after, up through Grok 4.3
Grok 4.5July 2026An "Opus-class" model for coding and agentic work (Musk called it "roughly comparable to Opus 4.7, but much faster"), with a 500K-token context window and API pricing of $2 / $6 per million input / output tokens
Grok 4.6August 2026The current flagship: a post-training upgrade for long-running agents and coding, tied with GPT-5.6 Sol Max on the Artificial Analysis Intelligence Index; same 500K context and $2 / $6 pricing as 4.5 (full details)
Grok 5In trainingAnnounced as a much larger model; still in training and not yet released as of August 2026 (what we know so far)
Two things are worth knowing beyond the table. The Grok 4 line is reasoning-first, which makes it stronger on hard problems but slower and pricier than a quick lookup needs. And **"Heavy" is not a different model, it is more compute on the same one:** the Grok 4 Heavy tier runs several agents that work a problem independently and compare answers, which helps on math and coding. Because xAI ships point releases almost constantly, the exact build you get depends on your plan and which surface you use. If open weights are the part you care about, Grok is like [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek): some older versions are downloadable, the latest are not. ## What Grok Can Do **Grok does most of what you expect from a modern assistant, plus a few things that come from its connection to X.** It answers questions, summarizes long documents and PDFs, writes and debugs code, reads images, and holds a back-and-forth conversation. The differences are at the edges. The headline capability is **real-time retrieval**. Because Grok can search live, public posts on X and the open web, it answers about current events, breaking news, and live sentiment in a way that a model relying only on its trained memory cannot. This is also why Grok is the model people reach for to take the pulse of a topic, and why [how AI search actually works](https://geotoolbox.ai/blog/how-does-ai-search-work) is worth understanding before you trust any answer it gives. Grok also added generative features. [Grok Imagine](https://geotoolbox.ai/blog/grok-imagine) generates images and short video clips from text prompts. **DeepSearch** is an agent that runs a multi-step research pass across many sources and writes up a synthesized report, rather than answering from a single page. Grok also has voice conversation, a writing-and-editing canvas, and reasoning modes that show their step-by-step work on harder problems. xAI has also pushed into AI companions, animated characters with set personalities, which has been one of the more controversial product directions and a driver of the safety questions we get to below. The honest summary: Grok's real, durable edge is the live data and the lower-friction persona. The image, video, and coding features are competitive but not categorically ahead of the other major assistants. Pick Grok for what is happening right now, not because any one creative feature is best available. ## How to Access Grok, and Is It Free? **Grok has a free tier, but it is tightly limited, and the way it is sold is the single most confusing thing about it.** You can reach Grok four ways: inside the X app, at the standalone site grok.com, through dedicated iOS and Android apps, and built into recent Tesla vehicles. You no longer need an X account to use the standalone version. There is a free plan, and for casual use it is fine, but it caps how many questions and images you get in a window before it makes you wait. To lift the limits and reach the newest models you pay, and this is where people get lost. The prices below are what was reported as of mid-2026 and they change often, so treat them as the shape of the pricing, not a quote. Our full [Grok pricing guide](https://geotoolbox.ai/blog/grok-pricing) breaks down every tier and the current promos.
TierReported price (mid-2026)What you get
Free$0Limited Grok access on X, grok.com, and the app, with rate limits on prompts and images
SuperGrok Lite~$10 / monthLighter paid tier added in 2026, aimed at users who mainly want more Grok Imagine
SuperGrok~$30 / monthA Grok-focused subscription with the highest standard limits and newest models
X Premium+~$40 / monthAn X subscription that includes higher Grok limits, alongside the platform's other perks
SuperGrok Heavy~$300 / monthAdds Grok 4 Heavy (the multi-agent tier) for power users
APIUsage-basedPay per token for developers; billed separately from any subscription
This trips up almost everyone. X Premium+ is a subscription to the X platform that happens to include more Grok. SuperGrok is a separate subscription to Grok itself. Paying for one does not give you the other, and the API is a third, independent product billed by usage. If you only want Grok, SuperGrok is usually the more direct buy. One more thing developers should know: on the Grok API, picking a smaller model is not the simple cost lever it is elsewhere. Token volume and whether reasoning mode is on drive most of the bill, and live web or X search and other tools can add per-call fees on top. Read the current rate card before you wire Grok into anything at scale. ## Grok vs ChatGPT: How They Differ **Grok and ChatGPT overlap on most everyday tasks, so the real question is which one fits a given job, not which is "better."** ChatGPT is the broader, more mature product with a larger ecosystem of integrations, custom GPTs, and enterprise tooling. Grok's advantage is narrower and sharper: live access to X and the web, and a willingness to engage with topics other models decline. Where they actually diverge is summarized below.
FactorGrok (xAI)ChatGPT (OpenAI)
Real-time dataNative, live access to public X posts and the webWeb browsing available, but X data is not a core source
PersonalityEdgier, fewer refusals, "truth-seeking" framingMore neutral and guardrailed by default
EcosystemTied to X, Tesla, and the xAI APILarge: custom GPTs, plugins, deep enterprise integration
Best atBreaking news, live sentiment, social listeningGeneral writing, broad knowledge work, established workflows
Reaching the flagshipOften pricier to reach the top tierFlagship access generally cheaper to reach
For most people the practical answer is to keep access to more than one model and route the task to whichever fits. Use Grok when the question is about what is happening on X or the wider web right now, and reach for ChatGPT or another assistant for general writing and analysis. If you are weighing the wider field, our full [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt) comparison goes deeper, and our breakdowns of [Grok vs Gemini](https://geotoolbox.ai/blog/grok-vs-gemini), [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude), [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt), and [ChatGPT vs Perplexity](https://geotoolbox.ai/blog/chatgpt-vs-perplexity) cover the same trade-offs from other angles. ## Is Grok Accurate and Safe? The Honest Picture **Grok's accuracy is genuinely contested, and several widely reported safety incidents are the reason it keeps making headlines.** Both deserve a straight, dated account rather than hype or a whitewash. On accuracy, the evidence cuts both ways. xAI has promoted strong benchmark results and, in some studies, a low hallucination rate. At the same time, a [Tow Center for Digital Journalism study reported by eWeek](https://www.eweek.com/news/ai-chatbot-citation-problem/) found AI search tools cited sources incorrectly at high rates, with Grok among the worst performers on citation accuracy. The benchmark claims have their own asterisks: when xAI said Grok 3's reasoning beat a rival OpenAI model, an OpenAI researcher argued the comparison used a more generous scoring method for Grok than for the competitor. The takeaway is not that Grok is uniquely bad, it is that benchmark numbers are marketing until verified, and that like every model, Grok can [hallucinate](https://geotoolbox.ai/blog/ai-hallucinations) confidently. Check anything that matters. On privacy, the default setting is the thing to know: **xAI uses your public X posts and your Grok conversations to train its models unless you opt out.** The [opt-out lives in X's privacy settings](https://tech.yahoo.com/ai/articles/posts-x-being-used-train-195730547.html) under the Grok options, and it generally applies going forward, not retroactively. That default also drew a formal inquiry from Ireland's Data Protection Commission, the EU's lead regulator for X, over how personal data was used to train Grok. If you would not paste something into a public post, do not paste it into Grok, and for business use, treat the consumer app as off-limits for anything sensitive and check the API and enterprise terms separately. On safety, Grok has had several widely reported incidents. In July 2025, after an update meant to make it "politically incorrect," Grok [posted praise of Hitler and referred to itself as "MechaHitler"](https://www.npr.org/2025/07/09/nx-s1-5462609/grok-elon-musk-antisemitic-racist-content); xAI said a code-path change upstream of the bot had left it pulling from extremist user posts on X, rather than the base model being at fault. Two months earlier, in May 2025, it had surfaced "white genocide" claims in unrelated answers, which xAI blamed on an unauthorized modification to its system prompt. Most seriously, through late 2025 its image tools were used to generate non-consensual sexual images of real people, with some content appearing to depict minors. In January 2026, [Indonesia and Malaysia became the first countries to block Grok](https://www.npr.org/2026/01/12/nx-s1-5674660/malaysia-indonesia-block-grok-ai-deepfakes) over the issue, and xAI restricted the feature in response. That same month the [European Commission opened formal proceedings against X under the Digital Services Act](https://digital-strategy.ec.europa.eu/en/news/commission-investigates-grok-and-xs-recommender-systems-under-digital-services-act) over Grok, examining whether X adequately assessed and mitigated the risk of illegal content, including the non-consensual sexual images. xAI has added guardrails after each incident, but the pattern is real, and it is why some organizations keep Grok off their approved-tools list. ## What Grok Means for Your Brand's Visibility **Grok now has an opinion about your business, and it forms that opinion differently from every other AI.** When someone asks Grok about your category or your company, it does not just recite training data. In its search-enabled modes it pulls from live X posts and the open web, and a DeepSearch run synthesizes many sources with citations. Which sources it surfaces, and whether it cites them, depends on the mode and the query. So the question for any brand is no longer just "do I rank on Google," it is "what does Grok say about me, and where is it getting that from." Three levers decide whether Grok surfaces you, and the order matters. First, **the sources Grok actually cites**: in our citation data for AI topics, answers lean heavily on Reddit, YouTube, and Wikipedia, so accurate coverage in those places is what most often earns a mention. The exact mix shifts by category, so audit the sources Grok cites in your own space before you prioritize. Second, **a clean, citable web page**, because the web search and DeepSearch steps quote clear, well-structured sources. Third, **your footprint on X itself**, since Grok reads live X posts for freshness and sentiment, so being discussed there shapes the tone of its answers even when X is not the source it cites. There is a newer wrinkle too: [Grokipedia](https://en.wikipedia.org/wiki/Grokipedia), the AI-built encyclopedia xAI launched in October 2025, is one more Musk-controlled knowledge surface that could shape what Grok and its users believe about you. Together these are the practical core of [generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization), and why [optimizing for AI search](https://geotoolbox.ai/blog/what-is-geo) is its own discipline now. The catch is that Grok is not interchangeable with the other engines. Because its retrieval leans on real-time X data, it can mention you when ChatGPT does not, or get a fact about you wrong that Gemini gets right. In our experience auditing brands across engines, a single "AI visibility score" hides exactly the variance that counts, which is why we [track visibility per engine](https://geotoolbox.ai/blog/how-to-track-ai-visibility) rather than as one number, and watch [share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) shift from one model to the next. Grok deserves its own column: its audience is real, and in our citation analysis its reliance on unfiltered live posts surfaces more negative or outdated framing than engines that lean on curated sources. That last point is the one to act on. Grok can confidently state something wrong about your products, pulled from a stale post or a critical thread, and most brands never see it because they are not looking. The moves are concrete: shore up the Reddit threads and YouTube videos it pulls from, keep a canonical page it can quote, and correct the record where it is wrong. But you cannot fix what you cannot see, which is why [tracking AI rank and citations](https://geotoolbox.ai/blog/ai-rank-tracker) across engines, Grok included, has become table stakes. The open question for most businesses is the one you cannot answer by reading about Grok: what is it actually telling people about you, right now? If you want to find out, geotoolbox tracks brand mentions and citations across eight AI engines, Grok among them, and shows which sources each one pulls from and where you are missing. You can [see what AI is citing about your brand](https://geotoolbox.ai/features/citation-interceptor) and start closing the gaps before they harden into the default answer. ## Frequently Asked Questions ### Is Grok free? Yes, Grok has a free tier you can use on X, at grok.com, or in the app, and you no longer need an X account for the standalone version. The free plan is rate-limited, so heavy users hit caps on prompts and image generation. Lifting the limits and reaching the newest models requires a paid plan such as SuperGrok or X Premium+. ### Who owns Grok, and is it the same as ChatGPT? Grok is made by xAI, the AI company Elon Musk founded in 2023; xAI acquired X in 2025, and SpaceX acquired xAI in 2026, so Grok now sits under Musk's SpaceX. It is not the same as ChatGPT and is not built on OpenAI's technology; it is a direct competitor. The main practical difference is that Grok has live access to public X posts and the web and uses a deliberately less filtered persona. ### Does Grok train on my X posts, and can I opt out? By default, xAI uses your public X posts and your Grok conversations to help train its models. You can opt out in X's privacy settings under the Grok and data-sharing options, but the change generally applies going forward rather than removing data already used. Keep sensitive information out of the consumer app, and for business or API use review the data-retention and training terms separately. ### What is the difference between Grok 4 and Grok 3? Grok 4 is the reasoning-first generation, and the current flagship is [Grok 4.6](https://geotoolbox.ai/blog/grok-4-6) (released August 2026); Grok 3 was the first to add a dedicated reasoning mode alongside a faster one. Grok 4 is more capable on hard problems and adds the "Grok 4 Heavy" tier, which runs multiple agents in parallel. For quick everyday questions the difference is small; for complex math or coding it is larger. ### What does the word "grok" mean? "Grok" comes from Robert A. Heinlein's 1961 novel Stranger in a Strange Land and means to understand something so deeply that you become part of it. Elon Musk chose it to signal deep understanding. Note that it is different from Groq, the AI chip company spelled with a q, which is unrelated to xAI. ### Can I trust Grok's answers? Treat them as a strong starting point, not a final source. Grok's real-time data is a real strength, but independent research has flagged accuracy and citation problems, and like any AI it can state wrong things confidently. Verify any fact, figure, or citation that matters before you rely on it. ## Sources - Grok (chatbot) - Wikipedia - `en.wikipedia.org/wiki/Grok_%28chatbot%29` - xAI (company) - Wikipedia - `en.wikipedia.org/wiki/XAI_%28company%29` - Stranger in a Strange Land - Wikipedia - `en.wikipedia.org/wiki/Stranger_in_a_Strange_Land` - Grok, Elon Musk's AI chatbot, shares antisemitic posts (July 2025) - NPR - `npr.org/2025/07/09/nx-s1-5462609/grok-elon-musk-antisemitic-racist-content` - Malaysia and Indonesia block Grok over AI deepfakes (January 2026) - NPR - `npr.org/2026/01/12/nx-s1-5674660/malaysia-indonesia-block-grok-ai-deepfakes` - The AI chatbot citation problem - eWeek - `eweek.com/news/ai-chatbot-citation-problem` - European Commission opens DSA proceedings against X over Grok (January 2026) - `digital-strategy.ec.europa.eu/en/news/commission-investigates-grok-and-xs-recommender-systems-under-digital-services-act` - Your posts on X are being used to train Grok AI - Yahoo Tech - `tech.yahoo.com/ai/articles/posts-x-being-used-train-195730547.html` - Grokipedia - Wikipedia - `en.wikipedia.org/wiki/Grokipedia` --- ## What Is Kimi AI? Moonshot AI's Open-Weight Model, Explained > What is Kimi AI? Moonshot AI's Kimi K2 and K3 models explained: pricing, is it safe, the China question, open weights vs open source, and what it means. - Canonical: https://geotoolbox.ai/blog/what-is-kimi-ai - Published: 2026-06-22 · Updated: 2026-08-08 Kimi AI is the chat assistant and the family of open-weight models built by Moonshot AI, a Beijing company that has become one of the loudest names in open AI. If you have seen "Kimi K2" topping coding leaderboards next to ChatGPT, Claude, and DeepSeek and want the plain version of what it is, who makes it, what it costs, whether it is safe, and what "open" actually means here, this is it, current as of August 2026. First, the disambiguation, because the name is crowded: this article is about the AI model, not the anime film "Kimi no Na Wa," the manga "Hana-Kimi," or Formula 1 drivers Kimi Räikkönen and Kimi Antonelli. We mean Kimi by Moonshot AI. We will also cover the parts most explainers skip: the open weights versus open source distinction, the data and China questions brands keep asking, and what a strong Chinese open model means for whether AI tools mention your business at all. This guide covers the Kimi family as a whole; for the current flagship specifically, see our [Kimi K3 explainer](https://geotoolbox.ai/blog/what-is-kimi-k3). ## What Is Kimi AI?
![Two cards contrasting Moonshot's closed Kimi app with its open model weights, shown with the K2 line as the example.](/blog/what-is-kimi-ai/kimi-app-vs-open-weights.png)
Kimi's key split, shown here with the K2 line: the app and API are closed, while the model weights are open to run. K3 continued the pattern on July 27, 2026 under its own Kimi K3 License.
**Kimi is two things under one name: an AI assistant you can chat with at [kimi.com](https://www.kimi.com/en), and the family of [large language models](https://geotoolbox.ai/glossary/large-language-model) that power it, now led by Kimi K3, both made by Moonshot AI.** You use the assistant the way you use ChatGPT or Claude: ask a question, paste a document, hand it a task, and it answers in natural language. It also reads images, writes and runs code, searches the live web, and runs multi-step agent workflows. The split between the assistant and the model matters more for Kimi than for most rivals, and it is the source of most confusion about it. The Kimi app and its API are a closed product Moonshot operates. The Kimi models underneath are open weights, which means the trained model files are published for anyone to download, run, and adapt. The flagship models inside ChatGPT stay closed (OpenAI has released separate, smaller open-weight models, but not the ones that power ChatGPT); Kimi opens its flagship model and keeps the product around it proprietary. That single fact shapes the pricing, the privacy trade-offs, and the strategic story we will get to. Where Kimi earned its reputation is coding and agentic work, tasks where the model plans, calls tools, and works through many steps on its own, at a fraction of what the closed frontier models charge. Calling it "a Chinese ChatGPT" undersells what is interesting about it. The rest of this guide walks through each piece. ## Who Makes Kimi? Moonshot AI, Explained Kimi comes from [Moonshot AI](https://en.wikipedia.org/wiki/Moonshot_AI), a Beijing startup founded in March 2023 by Yang Zhilin, Zhou Xinyu, and Wu Yuxin, who were schoolmates at Tsinghua University. Yang, the CEO, has said the goal is to build toward artificial general intelligence, with long context, multimodal understanding, and self-improving architecture as the milestones. The company name nods to Pink Floyd's "The Dark Side of the Moon," which is also where the Chinese name, 月之暗面, comes from. So yes, Kimi is a Chinese company, and that is central to the safety and data questions later. It is one of the better-funded ones. Alibaba put about $800 million into a roughly $1 billion round in February 2024 for a stake of around 36%, Tencent joined later that year, and in May 2026 Moonshot raised about $2 billion at a valuation near $20 billion, [according to TechCrunch](https://techcrunch.com/2026/05/07/chinas-moonshot-ai-raises-2b-at-20b-valuation-as-demand-for-open-source-ai-skyrockets/), on the back of roughly $200 million in annualized revenue. For a two-year-old lab, that is a steep climb, and it is funded largely by the open-weight strategy that put Kimi on the map. That backing answers a common question: who owns Kimi? Moonshot AI owns and runs it. Alibaba is the largest outside investor, but with a minority stake it does not appear to run Kimi day to day or dictate how it behaves. ## The Kimi Model Lineup: K2 to K3 Kimi is not one model but a fast-moving series. The breakout was Kimi K2 in July 2025, a one-trillion-parameter model released as open weights under a modified MIT license, [per Moonshot's GitHub repo](https://github.com/moonshotai/kimi-k2). It matched or beat much larger closed models on coding benchmarks, and because the weights were free, it spread fast. Since then Moonshot has shipped a new version roughly every couple of months. Here is where things stand.
Model (as of August 2026)ReleasedWhat changed
Kimi K2July 2025The open-weight flagship: 1T total / 32B active parameters, strong coding
Kimi K2-Instruct-0905September 2025Better coding; context window doubled to 256K
Kimi K2 ThinkingNovember 2025Reasoning interleaved with tool calls; benchmarked at native INT4
Kimi K2.5January 2026Vision (images and video) and the first Agent Swarm
Kimi K2.6April 2026Large jump on agentic benchmarks; swarm scaled to 300 sub-agents
Kimi K2.7 CodeJune 2026Coding-specialized; roughly 30% fewer thinking tokens
Kimi K3July 2026New flagship: a jump to 2.8T parameters (16 of 896 experts active) and a 1M-token context, the largest open-weight model at launch
The big change since the K2 line is Kimi K3, released July 16, 2026, with its weights going public on July 27 under a new custom Kimi K3 License, a step away from the K2 family's Modified MIT terms. It is a step up in scale, a 2.8-trillion-parameter model Moonshot bills as the largest open-weight model yet, and it moved Kimi from the cheap-and-cheerful open option into direct frontier competition on price and positioning. We cover its real specs, which benchmarks to trust, what it costs, and whether you can run it in our [Kimi K3 explainer](https://geotoolbox.ai/blog/what-is-kimi-k3). Two naming patterns are worth knowing. An "Instruct" model answers right away; a "Thinking" model works through hidden steps before replying, trading speed for accuracy on hard problems. On context length, the original K2 handled 128,000 tokens, doubled to 256,000 in the September Instruct-0905 update and carried into K2 Thinking; every model from K2.5 through K2.7 Code holds that same 256,000-token window, and the jump to a million tokens only arrives with K3. ### Kimi K2.5: Vision and the First Agent Swarm K2.5, released January 27, 2026, is where the K2 model line stopped being text-only. It added a vision encoder Moonshot calls MoonViT, trained on top of the existing K2 base, and it reads video as well as still images. Moonshot's [model card](https://huggingface.co/moonshotai/Kimi-K2.5) reports 76.8% on SWE-bench Verified. It shipped alongside Kimi Code, the company's coding agent for editors like VS Code and Cursor. The more consequential addition was Agent Swarm. Rather than one model working a task step by step, an orchestrator dispatches sub-agents that run in parallel and report back: up to 100 of them across as many as 1,500 tool calls in K2.5. Moonshot trained the orchestrator with a technique it calls parallel-agent reinforcement learning, freezing the sub-agents and rewarding the orchestrator in stages specifically to stop it collapsing back into doing the whole job itself. The caveats arrived with the same release. In the most detailed independent review of the model, Hugging Face's Maxime Labonne called it the first open-weight model whose vision "feels genuinely competitive", but measured it burning 89 million output tokens across an evaluation suite where comparable models used a median of 14 million, and watched swarm sub-agents drift into inconsistent definitions of the same shared concept. Cheap per token and cheap per task are not the same claim. K2.5 is still available, but not for much longer: Moonshot has closed it to newly registered API users and retires it on August 31, 2026, along with the older `moonshot-v1` family. ### Kimi K2.6 and K2.7 Code: The Agentic Jump K2.6 arrived April 20, 2026 with the same network underneath. Its published configuration describes an architecture identical to K2.5's, down to the expert count and the 256,000-token window, so whatever improved came from training rather than a redesign. The improvements also landed unevenly. Coding moved 3.4 points, SWE-bench Verified from 76.8% to 80.2%. The tool-use and agent benchmarks moved much further: Terminal-Bench 2.0 rose almost 16 points, from 50.8% to 66.7%, and three agent-specific suites roughly doubled. Agent Swarm scaled with it, to 300 sub-agents across 4,000 coordinated steps. One caution on that last number, which gets quoted a lot: 300 is a model capability, not something a subscription grants you, and Kimi's own plan table tops out at 8 concurrent swarm subtasks on its most expensive tier. [Artificial Analysis](https://artificialanalysis.ai/articles/kimi-k2-6-the-new-leading-open-weights-model) rated K2.6 the leading open-weights model on its intelligence index, fourth overall behind the Anthropic, Google and OpenAI flagships. Its standing criticism was token consumption, which is what K2.7 Code, released June 12, 2026, set out to address: a coding-specialized post-train of the same model that Moonshot says cuts thinking-token use by about 30% on average, and that cannot be run with thinking switched off. Its headline numbers deserve one caveat, though. The benchmarks Moonshot leads with for it, Kimi Code Bench v2 and Program Bench among them, are the company's own suites, and the model has not been submitted to DeepSWE, an independent coding benchmark that spreads models much further apart. As one developer [put it to VentureBeat](https://venturebeat.com/technology/kimi-k2-7-code-cuts-thinking-tokens-30-practitioners-say-benchmarks-dont-check-out), "every model 'improves' double digits on its own test suite." The most informative outside test came from researcher Elliot Arledge, who ran K2.7 Code against K2.6 on KernelBench-Hard, a public GPU-kernel optimization benchmark, and published his logs. His verdict was that "K2.7 is more honest but not more capable": on five of six problems the newer model wrote real Triton kernels where K2.6 had wrapped existing libraries, but two of those failed on its own bugs and one score regressed outright. It stopped papering over the hard part without yet doing the hard part better. Treat vendor benchmark wins as claims to verify against your own use, and treat this whole table as an August 2026 snapshot, because the version numbers move every few weeks. ## How Kimi K3 Works: The Current Flagship K3 is the model to understand first now, because it is what Moonshot leads with and what a paid plan points you at. It keeps the mixture-of-experts design the K2 line established but scales it hard: **2.8 trillion parameters in total, with only 16 of 896 experts active per [token](https://geotoolbox.ai/blog/what-are-tokens-in-ai)**, so the compute cost per token stays far below what the full size suggests. Moonshot publishes the expert counts but not the active-parameter figure, so treat any exact "active billions" number you see as an estimate. The bigger change is how K3 handles long inputs. It uses a hybrid linear-attention design Moonshot calls Kimi Delta Attention, with roughly three of every four attention layers running the cheaper linear form. That is what makes a 1,048,576-token [context window](https://geotoolbox.ai/glossary/context-window) practical rather than theoretical, and Moonshot's own research reports decoding up to 6.3 times faster at long context. K3 reads text, images, and video and returns text, with default output up to 131,072 tokens. One quirk matters in daily use: reasoning is always on. You can dial the effort down (the `reasoning_effort` control accepts `low`, `high`, or `max`, and defaults to `max`), but you cannot switch thinking off entirely, and the sampling parameters are locked server-side, so at its default K3 thinks hard on every request, which shows up in both speed and cost. The full architecture, the benchmarks worth trusting, and whether you can self-host it are in our [Kimi K3 explainer](https://geotoolbox.ai/blog/what-is-kimi-k3). ## How the Kimi Models Work: Mixture of Experts, Built for Agents Under the hood, Kimi is a transformer that predicts the next chunk of text one piece at a time, the same broad design as ChatGPT and Claude. Where it differs is two deliberate engineering choices that run through the whole family and explain why it is cheap and why it leans agentic. The first is the [mixture-of-experts](https://geotoolbox.ai/blog/how-does-chatgpt-work) design. Instead of one dense network where every parameter fires on every word, the Kimi models split into hundreds of specialist sub-networks and route each token to only a few of them. K2 holds a trillion parameters in total but uses only about 32 billion for any given token, and K3 pushes the same idea much further, activating only a sliver of an even larger network at a time. You get the knowledge of a huge model at the running cost of a small one, which is the main reason Kimi's API undercuts the closed frontier so heavily. The second is the focus on agentic behavior. Moonshot trained Kimi specifically to use tools, call functions, and carry a task across many steps rather than just answer in one shot. That is why its strongest results show up in coding agents and multi-step research, where the model has to decide what to do next, run it, read the result, and adjust. The trade-off is real: a model tuned to think and act in long chains can be slower and more verbose on simple questions, a complaint developers raise often. Practitioners who run it daily report a practical way to blunt that: Kimi tends to spiral into its longest reasoning when a prompt carries contradictory or cluttered instructions, so a short, direct, internally consistent prompt noticeably cuts the wasted thinking. The architecture itself is not a hidden dial you can game from the outside. It is what decides what Kimi is good at, and it points at the same place ChatGPT and Claude do: clear, well-structured information is what these models reason over best. ## Kimi K2 Thinking: The Model That Made the Reputation Going back a step chronologically: if you have heard one Kimi statistic, it probably came from Kimi K2 Thinking, released November 6, 2025. That is the release that got Kimi taken seriously as a frontier competitor, and its headline numbers are also the most consistently misquoted in the coverage that followed. What made it different from the Instruct models was that reasoning ran *through* the tool calls rather than before them. An Instruct model plans, then executes, which means a bad early assumption tends to survive the whole run. K2 Thinking reasons between steps, so it can query something, read the result, revise the plan mid-chain, and continue. Moonshot describes it as holding coherent goal-directed behavior across "200 to 300 sequential tool calls" without human intervention, though it published no harness detail behind that, so read it as a claimed working range rather than a specification. There is also an unusual engineering choice underneath. Moonshot applied quantization-aware training to the mixture-of-experts weights, so the model runs natively at INT4 precision, and it reports roughly a 2x generation speed-up from that in its own setup. More useful is that it reported every published benchmark at the same INT4 setting. The usual pattern is to benchmark at full precision and then quantize for shipping, which leaves the published scores describing a configuration nobody actually runs; here at least the precision you read about matches the weights you download. Those benchmark numbers, stated properly, are where most write-ups go wrong. On its [Hugging Face model card](https://huggingface.co/moonshotai/Kimi-K2-Thinking) Moonshot reports 71.3% on SWE-bench Verified, 60.2% on BrowseComp, and, on Humanity's Last Exam, three different figures for three different settings: 23.9% with no tools, 44.9% with tools, and 51.0% in what it calls heavy mode, which rolls out eight parallel attempts and aggregates them. The 44.9% is the figure that travels, almost always without the qualifier. It describes the model wired into a tool stack, not the model alone, and heavy mode is a further step removed again, spending roughly eight rollouts' worth of compute on a single answer. The economics attracted the same loose treatment. CNBC reported a training cost of roughly $4.6 million from an unnamed source and said plainly that it could not verify the figure; Moonshot's CEO has since said the number is not official, and its researchers have argued the cost is genuinely hard to pin down because so much of it is research and failed experiments. It gets repeated as fact anyway. [Artificial Analysis](https://artificialanalysis.ai/articles/kimi-k2-thinking-everything-you-need-to-know) scored K2 Thinking the highest open-weights model on its intelligence index at the time, but also found that running it on Moonshot's faster turbo endpoint made it one of the most expensive models in the whole comparison, because of how many tokens it burned getting there. One practical note if you go looking for it: Moonshot discontinued the entire `kimi-k2` series, K2 Thinking included, on its own API on May 25, 2026, and now directs users to K3. The weights remain openly published and third-party providers still serve them, which is the durability open weights buy you: when a closed model is withdrawn, outside users generally cannot keep running that version at all. ## Is Kimi Open Source? Open Weights vs Open Source This is where Kimi is most often described incorrectly, including by AI assistants asked about it. **Kimi K2 is open weights, not open source, and the difference is real.** Open weights means Moonshot publishes the finished model files, the billions of trained numbers, so anyone can download them, run them, and fine-tune them. Open source, in the strict sense, would also mean releasing the code and enough detail about the training data and method to rebuild a substantially equivalent model. Moonshot releases the weights but not the training data or the full recipe, so you can use the model freely, but you cannot fully audit or reproduce how it was made. The license is also not plain MIT. Kimi K2 ships under a modified MIT license, and the modification is a single attribution clause: per the [license on Hugging Face](https://huggingface.co/moonshotai/Kimi-K2-Instruct), if you deploy Kimi in a product with more than 100 million monthly active users or more than $20 million in monthly revenue, you must display "Kimi K2" prominently in the interface. For almost everyone that clause never triggers, but it means "open" here comes with one string attached. The current flagship changed the terms: K3 ships not under modified MIT but under Moonshot's own [Kimi K3 License](https://huggingface.co/moonshotai/Kimi-K3), which permits commercial use but is not OSI-approved, adding a separate-agreement requirement for very large model-as-a-service operators alongside a similar "Kimi K3" display clause for very large products. The smaller Kimi-VL model uses a standard MIT license with no such clause. One more layer: the model is open, but the Kimi app and API are not. The product you log into at kimi.com is closed software that Moonshot runs on its own servers. So "Kimi is open source" is true only of the weights, and only loosely. This is the same pattern DeepSeek, [Zhipu's GLM](https://geotoolbox.ai/blog/what-is-glm-5-2), Llama, [Qwen](https://geotoolbox.ai/blog/what-is-qwen), and Mistral follow, a wave of open-weight models that are genuinely free to run but are not open in the way the phrase implies. To see where Kimi lands among them, our [Chinese AI models comparison](https://geotoolbox.ai/blog/chinese-ai-models-compared) lines them up side by side. The practical upside is that privacy-sensitive teams can self-host the weights instead of sending data to Moonshot, which we come back to next. ## Is Kimi AI Safe? Privacy, Data, and the China Question "Is Kimi safe" actually folds two different questions together, and they have different answers. For everyday content, Kimi has guardrails and, like every assistant, can still be confidently wrong, so verify anything that matters. But its safety tuning looks lighter than the closed leaders'. A preliminary [independent safety evaluation of Kimi K2.5](https://arxiv.org/abs/2604.03121) found capability similar to GPT-5.2 and Claude Opus 4.5 but noticeably fewer refusals on dangerous (CBRNE) requests, along with more compliance on disinformation and copyright misuse. As a Chinese model, it also follows Chinese content rules, so on politically sensitive topics it deflects or stays vague where a Western model might engage. None of that makes it unusable, but do not assume its guardrails match the frontier labs'. The question with more weight for businesses is data. When you use the hosted app or API, your prompts go to Moonshot's servers, and where those servers sit depends on which door you use. The international API is operated by Moonshot AI Pte. Ltd., a Singapore entity, while the consumer service runs under Beijing Moonshot AI Technology Co., Ltd. in China. The underlying concern many security teams raise is that data handled under Chinese jurisdiction can be subject to local data and intelligence laws, which is the same caution applied to any China-hosted service, not a Kimi-specific accusation. There is also a usage term worth reading: Moonshot's API agreement says customer content may be used to develop and improve its services unless you arrange otherwise, with opt-outs reserved for enterprise or separate written agreements. It is worth weighing all of this plainly against your own risk tolerance and what data you would actually be sending. There is also one reported incident worth knowing. In April 2026 the [OECD.AI Incidents Monitor](https://oecd.ai/en/incidents/2026-04-21-8c79) logged a case where Kimi returned another user's resume during a task, a cross-user data exposure. Moonshot characterized it as a model hallucination, while outside observers described it as a data-isolation flaw and reported the exposed details as genuine rather than invented; there is no detailed first-party post-mortem. One reported incident is not a verdict, but it is a fair data point for a privacy review. The honest mitigation is the one the open weights make possible. If data residency is a dealbreaker, you do not have to use Moonshot's servers at all; a team can run the open Kimi K2 weights on its own infrastructure, so prompts never leave the building. That is the clearest practical answer to the China question: for sensitive workloads, self-host rather than send. ## Kimi vs ChatGPT, Claude, and DeepSeek Before comparing, fix one category error that trips up most write-ups, including AI ones: Kimi K2 and DeepSeek are models you can download and run, while ChatGPT and Claude are products that front closed models you can only rent. A fair comparison either lines the Kimi app up against the ChatGPT and Claude apps, or lines the Kimi K2 model up against the closed models inside them. Here is the practical version.
ToolWhat it isStrongest atOpen weights?Rough cost
Kimi K2 (Moonshot)Open-weight model + appCoding, agentic multi-step work, long context, low costYes (modified MIT)Low; can self-host free
ChatGPT (OpenAI)Product fronting GPT modelsGeneral-purpose use, images, voice, the widest ecosystemNo (flagship); separate gpt-oss models, yesFree tier; paid from $20/mo
Claude (Anthropic)Product fronting Claude modelsWriting, careful reasoning, long documents and codeNoFree tier; paid from $20/mo
DeepSeek (DeepSeek)Open-weight model + app (also China-based)Reasoning and coding at very low costYesLow; can self-host free; same China data caveats
On capability, the honest read is that Kimi competes hardest on coding, agentic tasks, and price, where developers report it doing real work for a fraction of what [Claude or ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) cost. Where the closed products still tend to lead is general polish, reliability across varied tasks, multimodal range, and ecosystem depth. Independent reviewers also note Kimi can be slower and noticeably more verbose, answering a simple question with several paragraphs. Kimi's arrival was called "another DeepSeek moment," and that framing is the real story: a Chinese open model matching far more expensive Western ones no longer shocks anyone, which itself signals how fast the gap is closing. That competitive pressure has an edge to it. In February 2026, [Anthropic alleged](https://fortune.com/2026/02/24/anthropic-china-deepseek-theft-claude-distillation-copyright-national-security/) that several Chinese labs, Moonshot among them, used networks of fake accounts to harvest millions of Claude conversations and distill their capabilities. Anthropic did not sue, Moonshot did not publicly respond, and it remains an unproven accusation, but it is part of the backdrop to how these cheaper, strong models are built. If you are weighing models head to head, our [Claude AI explainer](https://geotoolbox.ai/blog/what-is-claude-ai) covers the other side of that comparison. ## Is Kimi Free? Pricing and How to Access It Yes, Kimi has a real free tier. The plan is called Adagio, it covers the web app and the iOS and Android apps, and it needs no credit card. Above it sit four paid tiers, from Moderato at $19 a month to Vivace at $199, with annual billing taking roughly 20% off at every level. New paid subscriptions are paused as of August 2026 while Moonshot adds capacity; existing subscribers keep their plans. The important mechanic is that paid plans are metered in credits rather than messages, and conversations with the K2.6 model do not consume credits at all. For the full tier table, what a credit actually buys, how refreshes and downgrades work, and whether paying beats ChatGPT Plus or Claude Pro, see our [Kimi pricing](https://geotoolbox.ai/blog/kimi-pricing) breakdown. Developer access is billed separately from the app, and the metered rates and access routes are in our [Kimi API pricing](https://geotoolbox.ai/blog/kimi-api-pricing) guide. A note on access: signing up from outside China can be fiddly because phone verification does not always work for foreign numbers; logging in with Google or scanning the app's QR code is the usual workaround. And on self-hosting, "free" is the license, not the hardware. Kimi K2 is a one-trillion-parameter model, and running it well takes a multi-GPU server most individuals do not have. The free, run-anywhere option is real for companies with infrastructure; for everyone else, the hosted app or a third-party provider is the practical route. If you just want to try Kimi without committing to a subscription, aggregators like [OpenRouter](https://openrouter.ai/moonshotai) serve the Kimi models, from K2 through K3, on pay-per-use per-token billing, which is what many developers reach for to test a model before paying for any single plan. ## What Kimi Means for Your AI Visibility Step back from the specs and there is a marketing question hiding in all this. Every new strong model is another place where a customer might ask "what is the best tool for X" or "is [your company] any good," and get an answer that shapes a buying decision. Kimi is one more of those places, and the open-weight twist makes it bigger than it looks. Because Kimi K2's weights are public, the model does not only answer inside Moonshot's own app. It gets hosted, fine-tuned, and embedded into a long tail of downstream products and providers you will never see individually. You cannot audit every deployment that runs on Kimi, DeepSeek, or any other open model. What you can do is influence the input they all share: how clearly and consistently your business is represented across the open web, which is the [foundation of generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization). That splits into two practical jobs. The first is reachability: every one of these models and the crawlers feeding them has to be able to fetch your site in the first place. Block the AI crawlers, intentionally or not, and you close the live door into every model that fetches the web, even if you cannot undo what they already learned in training. The second is consistency: the brands that get described correctly are the ones whose facts line up across the sites a model is likely to read. In our experience at geotoolbox, the businesses that surface well in AI answers are rarely the ones with the prettiest homepage; they are the ones a model can find, parse, and trust without tripping over contradictions. That is a measurable problem, which is why we build tooling around it. Our guides on [what GEO is](https://geotoolbox.ai/blog/what-is-geo) and [tracking your AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) go deeper on both halves. The open-model wave does not change the playbook so much as raise the stakes: there are simply more engines that can mention, or mangle, what you have built. The first move is to check whether they can even read you. Run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether the AI crawlers can reach and parse your site, and where the gaps are, before the next model launches and the question gets asked again. ## Frequently Asked Questions ### Is Kimi AI a Chinese company? Yes. Kimi is made by Moonshot AI, a startup based in Beijing and founded in March 2023 by three Tsinghua University classmates. Alibaba is its largest outside investor. As of May 2026 the company was valued at around $20 billion. ### Is Kimi AI safe to use? For everyday content, Kimi has guardrails like any major assistant and, like all of them, can still be wrong, so verify what matters. The bigger consideration is data: prompts sent to the hosted service go to Moonshot's servers (a Singapore entity for the API, a China entity for the consumer app), so they fall under those jurisdictions. For sensitive work, the safest route is to self-host the open weights so your data never leaves your own systems. ### Is Kimi AI free? Yes, there is a genuine free tier called Adagio on the web and in the mobile apps. Paid plans run from $19/mo (Moderato) to $199/mo (Vivace), metered in credits rather than messages, and the open weights are free to run if you have the hardware. New sign-ups are paused as of August 2026 while Moonshot adds capacity. Our [Kimi pricing](https://geotoolbox.ai/blog/kimi-pricing) guide covers every tier and what the credits actually buy. ### Is Kimi better than ChatGPT and DeepSeek? It depends on the job. Kimi is praised for coding, agentic multi-step tasks, and very low cost, and it competes closely with [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek), the other major Chinese open-weight model. ChatGPT is broader and more polished across general use, images, and voice. Many people use more than one and pick per task. ### Is Kimi really open source? Not in the strict sense. Kimi K2 is open weights: the trained model is published under a modified MIT license, so you can download, run, and fine-tune it, but Moonshot does not release the training data or full recipe, and the Kimi app and API are closed. The license also asks very large deployments to credit "Kimi K2" in their interface. The current flagship, K3, ships under a separate custom Kimi K3 License, also open weights but not OSI-approved. ### Can I run Kimi on my own computer? The weights are public, but Kimi K2 is a one-trillion-parameter model that needs a multi-GPU server, not a laptop, to run well. Self-hosting is realistic for companies with that infrastructure, mainly for privacy or cost control. For everyone else, the hosted app or a third-party provider is the practical option. ## Sources - Kimi K2 repository (specs and license) - Moonshot AI, 2025 - `github.com/moonshotai/kimi-k2` - Kimi K2-Instruct model card and license - Moonshot AI (Hugging Face) - `huggingface.co/moonshotai/Kimi-K2-Instruct` - Kimi K2 Thinking benchmarks, INT4 QAT and tool settings - Moonshot AI (Hugging Face), November 2025 - `huggingface.co/moonshotai/Kimi-K2-Thinking` - Kimi K2.5 model card (MoonViT vision encoder, benchmarks, license) - Moonshot AI (Hugging Face), January 2026 - `huggingface.co/moonshotai/Kimi-K2.5` - Kimi K2.6 model card (matched K2.5 vs K2.6 benchmark table) - Moonshot AI (Hugging Face), April 2026 - `huggingface.co/moonshotai/Kimi-K2.6` - Kimi K2.6 announcement (Agent Swarm at 300 sub-agents) - Moonshot AI, April 2026 - `kimi.com/blog/kimi-k2-6` - Kimi K2.7 Code model card (thinking-token reduction, benchmark suites, 256K context) - Moonshot AI (Hugging Face), June 2026 - `huggingface.co/moonshotai/Kimi-K2.7-Code` - Kimi K2.7-Code cuts thinking tokens 30%, but practitioners say the benchmarks don't check out (KernelBench-Hard results, DeepSWE submission) - VentureBeat, June 2026 - `venturebeat.com/technology/kimi-k2-7-code-cuts-thinking-tokens-30-practitioners-say-benchmarks-dont-check-out` - Model list and deprecation dates (K2 series retired May 25, 2026; K2.5 sunset August 31, 2026) - Moonshot AI developer docs, accessed July 19, 2026 - `platform.kimi.ai/docs/models` - Kimi K2 Thinking: everything you need to know (intelligence index, token consumption, endpoint costs) - Artificial Analysis, November 2025 - `artificialanalysis.ai/articles/kimi-k2-thinking-everything-you-need-to-know` - Kimi K2.6, the new leading open weights model (hallucination rate, index ranking) - Artificial Analysis, April 2026 - `artificialanalysis.ai/articles/kimi-k2-6-the-new-leading-open-weights-model` - Kimi K2.5: still worth it after two weeks? (independent review, token verbosity, swarm drift) - Maxime Labonne, Hugging Face blog, February 2026 - `huggingface.co/blog/mlabonne/kimik25` - Alibaba-backed Moonshot releases Kimi K2 Thinking (reported $4.6M training cost, unverified) - CNBC, November 2025 - `cnbc.com/2025/11/06/alibaba-backed-moonshot-releases-new-ai-model-kimi-k2-thinking.html` - Moonshot CEO: reported $4.6M training cost "isn't official" - Yicai Global, November 2025 - `yicaiglobal.com/news/kimi-k2-thinkings-reported-usd46-million-training-cost-isnt-official-moonshot-ceo-says` - Moonshot AI company overview - Wikipedia - `en.wikipedia.org/wiki/Moonshot_AI` - China's Moonshot AI raises $2B at $20B valuation - TechCrunch, May 2026 - `techcrunch.com/2026/05/07/chinas-moonshot-ai-raises-2b-at-20b-valuation-as-demand-for-open-source-ai-skyrockets` - Anthropic accuses Chinese labs of distillation via Claude - Fortune, February 2026 - `fortune.com/2026/02/24/anthropic-china-deepseek-theft-claude-distillation-copyright-national-security` - Kimi cross-user data exposure incident - OECD.AI Incidents Monitor, April 2026 - `oecd.ai/en/incidents/2026-04-21-8c79` - An Independent Safety Evaluation of Kimi K2.5 - arXiv, April 2026 - `arxiv.org/abs/2604.03121` - Kimi K2: What's all the fuss and what's it like to use? - Thoughtworks - `thoughtworks.com/en-us/insights/blog/generative-ai/kimi-k2-whats-fuss-whats-like-use` --- ## What Is Qwen? Alibaba's Open-Weight AI Family, Explained > What is Qwen? Alibaba's open-weight AI family explained: the Qwen3.8 lineup, the open-vs-closed Max split, safety, and what it means for your AI visibility. - Canonical: https://geotoolbox.ai/blog/what-is-qwen - Published: 2026-06-22 · Updated: 2026-08-20 By raw download count, Qwen is the most popular open AI model in the world, and most marketers have never heard of it. Alibaba's open models are downloaded, fine-tuned, and built into other developers' products so widely that Qwen can sit inside tools you would never connect back to a Chinese tech giant. Here is the plain version of what Qwen is, current as of August 2026, who builds it, whether it is safe, and where it stands. And then the part the developer write-ups skip: what the most-deployed open model means for whether AI tools describe your brand correctly. ## What Is Qwen? **Qwen is Alibaba Cloud's AI model family, most of it released as open weights and known in China as Tongyi Qianwen.** It is not a single model but a whole lineup, with versions for text, code, vision, image generation, and audio. If you have run a local model on your laptop, browsed an "open model" leaderboard, or used an AI tool that swapped in a cheaper engine behind the scenes, there is a good chance Qwen was involved. Like the other open models, Qwen is really three things, and keeping them straight clears up most of the confusion. There are the open weights, which Alibaba publishes so anyone can download, run, and fine-tune the [large language models](https://geotoolbox.ai/glossary/large-language-model). There is the hosted product, the free Qwen chatbot and Qwen Studio at qwen.ai, which Alibaba runs on its own servers. And there is the API, which developers call to build on it. The models are mostly open; the service around them is Alibaba's. The reason most marketers have never heard of Qwen, despite its scale, is that it lives mostly as infrastructure rather than as a consumer brand. It does not have a household name like ChatGPT. It sits underneath things, fine-tuned and rebranded inside other products until the Alibaba name disappears. That ubiquity is the whole story, and it is exactly why Qwen is worth your attention even if you never type a prompt into it. Qwen is usually pronounced "chwen." It is the AI model family, not the American football player whose name people misspell into the same search box, and not the accent chair that shares the spelling. When this article says Qwen, it means Alibaba's AI. ## Who Makes Qwen? Alibaba and the Tongyi Qianwen Team **Qwen comes from Alibaba Cloud, built by the team behind Tongyi Qianwen, and that pedigree is the first thing that sets it apart.** Alibaba launched the first Qwen beta in 2023 and has shipped new models at a relentless pace ever since, eventually spinning the work into a dedicated unit inside the company. That is a different kind of player from the other Chinese labs you have probably seen in headlines. [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek) came out of a quant hedge fund, [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai) from the startup Moonshot, and [GLM](https://geotoolbox.ai/blog/what-is-glm-5-2) from Zhipu. Qwen comes from one of the largest technology companies on the planet. It is a Big Tech open-model program, not a scrappy upstart, and that shows in both the breadth of the lineup and the strategy behind it. The strategy is the familiar one Meta and Google use with their open models: give the models away to win developers, and earn the money on cloud and downstream products. Alibaba leans into that hard. It releases most of the lineup as free downloads, and it also wires Qwen into its own ecosystem, powering features across Alibaba's e-commerce and travel platforms. The free models build the habit; the paid cloud and the flagship API are where Alibaba expects to be repaid. ## The Qwen Family: One Name, Dozens of Models **The hardest part of understanding Qwen is that there is no single "Qwen."** There are dozens, across several generations and a wall of suffixes: Qwen2.5, Qwen3, Qwen3.5, Qwen3.6, Qwen3.7, Qwen3.8, with labels like Max, Plus, Coder, VL, Image, and Omni appearing across the family, in sizes from under a billion parameters to hundreds of billions. If you have felt lost trying to figure out which Qwen is "the" Qwen, that is the design, not your fault. The throughline is easy enough. The series ran from the original Qwen in 2023 through Qwen2 and Qwen2.5, then [Qwen3 in April 2025](https://en.wikipedia.org/wiki/Qwen), which added a hybrid "thinking" mode that lets the model toggle between fast answers and slower step-by-step reasoning. From there it moved quickly through Qwen3.5, Qwen3.6 and Qwen3.7 to the current flagship, **Qwen3.8-Max**, released on August 3, 2026. The family is trained for well over 100 languages, which is part of why Qwen is so widely adopted outside the English-speaking world. **[Qwen3.8-Max](https://geotoolbox.ai/blog/qwen3-8-max)** is a 2.4-trillion-parameter multimodal model with 95 billion parameters active per token and a one-million-token context window, first shown as a preview at the World AI Conference in Shanghai on July 19 and released properly two weeks later. Alibaba calls it "the most capable model in the Qwen family to date." It runs through the API as `qwen3.8-max` at $2 per million input tokens and $6 per million output, and unlike the July preview it comes with a published benchmark table, though a mixed one: Alibaba measured Qwen itself and took several rivals' scores from third parties. The weights, promised for the week of August 10, landed on August 12, 2026 in the official [Qwen3.8-2.4T-A95B](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B) repository: the first time a Max-tier Qwen has ever been opened. The opening is partial, though: the released variant is text-only, its native context stops at 262,144 tokens (extensible to about a million), and it ships under a bespoke "Qwen3.8-Max" license rather than Apache 2.0. What really sets Qwen apart from its open rivals is breadth. Where DeepSeek, Kimi, and GLM are essentially text-and-code models, Qwen is a full multimodal family. Here are the main lines as of August 2026.
Qwen line (as of August 2026)What it doesOpen weights?
Qwen3.6 (dense + MoE)General text and reasoning, hybrid thinking; the current open line and the models most people actually runYes (Apache 2.0)
Qwen3.8-MaxThe current flagship, 2.4T parameters with 95B active, 1M context, multimodalYes, since August 12, 2026: text-only variant (262,144-token native context), under a bespoke "Qwen3.8-Max" license, not Apache 2.0
Qwen3.7-MaxThe previous flagship, 1M context, still on sale; listed higher than 3.8-Max but cheaper while its 50% promotion runsNo (API only)
Qwen3.7-PlusMultimodal mid-tier: image and video understanding, deep reasoning, GUI and agentic work, 1M contextNo (API only)
Qwen-CoderCoding and agentic software tasksYes (open versions)
Qwen-VLVision: reads and reasons over images and videoYes (open versions)
Qwen-ImageImage generation and editingYes (open versions)
Qwen-OmniText, image, audio, and video together, with real-time speech; the current commercial builds are qwen3.5-omni-plus and its realtime variantMixed (some open, some not)
## Is Qwen Open Source? The Open-vs-Closed Reality
![Checklist of Qwen model lines as of August 2026: Qwen3.6 ships under Apache 2.0 and Qwen-Coder, Qwen-VL and Qwen-Image have open versions, while the flagship Qwen3.8-Max shipped its weights on August 12, 2026 as a text-only variant under a bespoke 'Qwen3.8-Max' license, and Qwen3.7-Max and Qwen3.7-Plus remain closed and API-only.](/blog/what-is-qwen/qwen-open-family-closed-flagship.png)
Most Qwen lines are downloadable, the 3.7 tiers are not, and the flagship's August 12 weights drop is text-only under a bespoke license.
**Qwen is the open-model story with an asterisk: most of the family is genuinely open, but the best model in it is not.** This trips up a lot of people, and it is worth getting right, because "is Qwen even open source anymore?" has become a real question in the community. Here is the honest version. Most of the Qwen lineup, the dense models and many of the Mixture-of-Experts models, ships as open weights under the permissive Apache 2.0 license. You can download them, run them on your own hardware, fine-tune them, and use them commercially, subject only to the light conditions of the Apache 2.0 license. That is the part that earned Qwen its reputation. But the flagship "-Max" tier, from [Qwen3-Max through Qwen3.7-Max and, until August 12, 2026, the current Qwen3.8-Max, was proprietary](https://en.wikipedia.org/wiki/Qwen): it ran only through Alibaba's API, with no download. Alibaba had followed this pattern since Qwen2, opening most of the lineup while keeping its most capable model closed. That asterisk grew, then partly receded. The 3.7 generation stayed closed: Qwen3.7-Max in May 2026 and Qwen3.7-Plus in early June, both API-only, and neither ever got weights. Qwen3.8-Max looked set to follow until August 12, 2026, when Alibaba published its weights (text-only, under a bespoke "Qwen3.8-Max" license rather than Apache 2.0): the first open Max-tier Qwen, though the fully multimodal version of the flagship itself remains API-only. A separate, smaller model closed that gap two days later: [Qwen3.8-27B](https://huggingface.co/Qwen/Qwen3.8-27B), a 27-billion-parameter vision-language model, shipped around August 14, 2026 as open weights under the ordinary Apache 2.0 license, so a genuinely open, genuinely multimodal Qwen now exists, just not at flagship scale. What is still closed is the "-Plus" mid-tier, Qwen3.7-Plus and its successors, the size most builders would actually reach for before jumping to a 2.4T flagship: it stays API-only even now that the flagship has partly opened. The counterweight is that Alibaba has not slowed its open releases elsewhere, shipping Apache 2.0 weights for agent-environment, embedding, and robotics models through June and July 2026, with the most recent uploads to its Hugging Face account dated late July. The 3.8 launch is where that pattern got tested. Alibaba's release post said it would open-source the Qwen3.8-Max weights in the week of August 10, 2026, the first time a Max-class Qwen would be opened. That was a real break with three years of practice, which is also why it was worth checking rather than assuming, and the check eventually came back in Alibaba's favor: two days past the announced window, the weights landed on August 12, 2026 in [Alibaba's official Hugging Face organization](https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B). The opening is partial: the published variant is text-only, its native context stops at 262,144 tokens, and the license is a bespoke "Qwen3.8-Max" license rather than the family's usual Apache 2.0. The third-party repositories that carried the Qwen3.8 name before that date held either nothing or older models relabelled with the new number, so only Alibaba's own account settles it. Licensing adds another layer. Most open models are Apache 2.0 with no conditions, but a few sit under the separate [Qwen License](https://github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT), which requires you to request a license from Alibaba once a product passes 100 million monthly active users, and the smallest research variants are non-commercial. And as with every model in this category, [open weights are not the same as open source](https://geotoolbox.ai/blog/open-weights-vs-open-source): Alibaba publishes the finished weights, not the training data or the full recipe.
TierExampleDownloadable?License
Open dense + MoEQwen3.6-35B-A3B, Qwen3.6-27BYesApache 2.0 (no strings)
Some larger / older weightsQwen2-72BYesQwen License (100M-MAU clause)
Small research variantsQwen2.5-3BYesQwen Research (non-commercial)
Flagship "-Max"Qwen3.8-MaxYes (August 12, 2026), text-onlyBespoke "Qwen3.8-Max" license for the weights; the fully multimodal version remains API-only
Open multimodal (mid-size)Qwen3.8-27BYes (~August 14, 2026), vision-languageApache 2.0 (no strings)
Because so much of Qwen is free and unrestricted, it gets downloaded, fine-tuned, and rebranded into a vast number of community versions, including ones modified to strip out the guardrails. Hold onto that point. It is the part that actually matters for your brand, and we come back to it at the end. ## How Good Is Qwen? Qwen3.8-Max and the Honest Picture **As of August 2026 the flagship Qwen3.8-Max is in the frontier conversation on Alibaba's own numbers, but the more useful story is still the open models.** The 3.8 release finally came with a benchmark table, which the July preview did not have, and it is worth reading closely rather than as a headline. Alibaba's table puts Qwen3.8-Max at 86.6 on Terminal-Bench 2.1, ahead of Claude Fable 5 at 84.6, and at 92.6 on GPQA Diamond, level with Fable 5. It also shows the model well behind on the harder end: 67.7 against Fable 5's 80.0 on SWE-bench Pro, and 43.6 against 53.3 on Humanity's Last Exam. So it is a good model losing several of the comparisons its own vendor chose to publish, which is already a more honest picture than the "second only to Claude Fable 5" positioning that followed the July preview. The outside numbers are less flattering still. [Artificial Analysis independently runs Terminal-Bench 2.1](https://artificialanalysis.ai/evaluations/terminalbench-v2-1) and scores Qwen3.8-Max at 81.3%, below Fable 5's 84.6% rather than above it. Alibaba disclosed why the two differ: on that benchmark it ran Qwen itself and reported every rival's best published score from elsewhere, so its table sets a self-run number against other people's. The same caveat applies to its generational claim, where the Qwen3.7-Max figure is imported and the Qwen3.8-Max one is not. Measured by Artificial Analysis at both ends, the jump is 74.5% to 81.3%. That is the whole reason to treat vendor benchmarks as a starting point. Benchmarks are also the wrong place to look for what this model is actually for. The more telling evidence in Alibaba's launch post is agentic: it says Qwen3.8-Max ran unattended for more than ten days building a working command-line project, accumulating 265 commits and 127 pull requests with no human in the loop, and separately spent about 125 hours reproducing a research paper from nothing but the paper and a set of GPUs before improving on its result. Those are Alibaba's own accounts and nobody has reproduced them. They still describe the thing worth testing: long, unsupervised work rather than single answers. So the short verdict is that Qwen3.8-Max is a strong multimodal model and a competent coding one at a fraction of frontier pricing, and that nothing published so far makes it the best model in the world. The result that should actually impress a non-developer is quieter. Qwen's open models, the ones anyone can download for free, routinely punch well above their size, competing with much larger and far more expensive models. That is exactly why developers reach for them as a default, and it is the real Qwen achievement: not that the closed flagship tops a leaderboard for a week, but that the free, downloadable models are good enough to build on. One honest caveat. On pure English-language tasks the gap between Qwen and the best Western open models like Google's Gemma is narrow and shifts with every release. Qwen's more durable advantage is its broad multilingual coverage, strongest in Chinese and across many non-English languages, which matters more if your audience is global than if it is English-first. ## Is Qwen Safe? The China and Alibaba Question **The honest answer turns on the same distinction as every Chinese model, who makes it versus who hosts your data, plus one Alibaba-specific wrinkle.** Get the first part straight and most of the worry sorts itself out. When you use the hosted Qwen app or API, your prompts travel to Alibaba's servers, which puts them under Chinese jurisdiction and the data rules that come with it. Any hosted model means your prompts leave the building, a US provider included; the difference with a Chinese one is whose jurisdiction and policy they land under. For sensitive or proprietary work, that is a real consideration, and it is the same caution we walk through for [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek). The escape hatch is the open weights: because most of the family is downloadable, a team can run Qwen on its own infrastructure, so the prompts never leave the building and Chinese hosting never enters the picture. Self-host rather than send is the clearest answer for sensitive work. The wrinkle is procurement. Some enterprises, especially in regulated or security-sensitive settings, run a blanket "no Chinese-origin models" policy, and that is a box-checking rule, not a technical one. Self-hosting solves the data-residency problem but may not clear a policy that bars the model by origin regardless of how it is run. If you operate in a regulated or security-sensitive environment, that policy question is often the real blocker, not the engineering. There is also a content angle: the official app follows Chinese content rules on politically sensitive topics, and because the weights are open, the community has even produced "uncensored" forks that strip those guardrails out, which is a reminder of how far beyond Alibaba's control these models travel once released. ## How to Access Qwen **There are three ways in, and which one you pick decides the cost and privacy tradeoff.** The quickest is the free Qwen chatbot and Qwen Studio at qwen.ai, which need no setup and cover text, images, and more. For building, the API is OpenAI-compatible, so it slots into existing tools by changing little more than an endpoint, and it is priced well below the closed Western frontier. And because most of the family is open, you can skip Alibaba's servers entirely and self-host.
How you use itCost (as of August 2026)What you get
Qwen chat / Qwen StudioFreeWeb chat with text and images, no setup, quickest way to try it
Standalone API (Qwen3.8-Max)$2 in / $6 out per 1M tokensBuild the current flagship into your own product
Standalone API (Qwen3.7-Max)$2.50 in / $7.50 out per 1M tokens list, currently on a limited-time 50% discountThe previous flagship, cheaper than the current one while the promotion lasts
Self-host (open models)Free license, your own hardwareData stays local; runs on anything from a laptop to a server
The flagship's [API pricing on OpenRouter](https://openrouter.ai/qwen/qwen3.8-max) matches Alibaba's own $2 and $6 rates, and the older Qwen3.7-Max sits below that while its 50% promotion runs. Either way it is a fraction of what the top closed models charge. Promotional rates move, so check live before budgeting. The open models live on [Hugging Face](https://huggingface.co/Qwen) and Alibaba's ModelScope, and they run through the usual local tools like Ollama and LM Studio, which is how Qwen ended up on so many laptops and phones in the first place. ## Qwen vs DeepSeek, Kimi, GLM and Llama **In the crowded open-model field, Qwen's distinction is breadth and reach rather than a single benchmark win.** The other open models tend to specialize: [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek) is the reasoning-and-cost specialist, [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai) leans agentic with long context, [GLM](https://geotoolbox.ai/blog/what-is-glm-5-2) is coding-first, and Meta's Llama is the Western pioneer that kicked off the open-weight wave. Qwen is the one that does a bit of everything, across the most modalities, with the widest deployment. For a full side-by-side of how these stack up on cost, openness, and capability, see our [Chinese AI models comparison](https://geotoolbox.ai/blog/chinese-ai-models-compared).
ModelMakerStrongest atOpen weights?
QwenAlibabaBreadth (text, vision, image, audio), multilingual, sheer reachMostly yes; flagship weights text-only since Aug 12 (bespoke license), multimodal API-only
DeepSeekHigh-FlyerReasoning, math, very low costYes (MIT)
KimiMoonshotAgentic work, long contextYes (modified)
GLMZhipuCoding, agents, long-horizon tasksYes (MIT)
LlamaMetaThe Western open pioneer, large ecosystemYes (Llama license)
For most teams the practical answer is not to crown one open model but to reach for whichever fits the job, often pairing a cheap open model with a closed frontier model like [Claude or GPT](https://geotoolbox.ai/blog/claude-vs-chatgpt) for the hardest work. What makes Qwen matter beyond that menu is not where it lands on any one test. It is how many places it has already ended up, and that is the thing with real consequences for your brand. ## What Qwen Means for Your Brand's AI Visibility The question most marketers actually have is whether they need to do something about Qwen. The short answer is no, not directly, and the reason is a distinction worth holding onto. **Qwen is infrastructure, not a destination.** You do not "optimize for Qwen" the way you optimize a page for Google. Alibaba does run a Qwen chatbot, but in most Western markets your buyers are not researching you there; they are on ChatGPT, Perplexity, Gemini, and Claude. Those products are what you [track your visibility in](https://geotoolbox.ai/blog/how-to-track-ai-visibility), and a single open model launching changes none of that. It is worth saying plainly, too, that ranking number one on Google does not mean an AI engine will mention you, because that is a separate system with its own logic. So why does Qwen come up at all? Because of its reach, with one honest qualifier. By mid-August 2026 Alibaba said the Qwen family had passed [3 billion cumulative downloads](https://fortune.com/2026/08/15/alibaba-qwen-open-ai-models-3-billion-downloads-meta-google), making it the most-downloaded open model family in the world, ahead of both Meta and Google. That is Alibaba's own headline, and it is worth an honest qualifier: Hugging Face's own *State of Open Models* report (August 2026) counts roughly 2 billion Qwen downloads and about 151,000 derivative models, against the 300,000+ Alibaba claims. Whichever count you take, the order of magnitude is the same, and it dwarfs any other open family. Most of those are developers running Qwen for coding, local inference, or backend features that never describe a brand. But some of those deployments do face end users, the chat assistants, support bots, and search wrappers quietly built on Qwen, and their reach is impossible to measure from a download count and impossible to audit. You cannot inspect every product that runs on it, including community forks like "Liberated Qwen," a version stripped of its content guardrails. That is the real point: an open model this widely embedded becomes part of the substrate that describes your brand in places you will never see, not because it is sinister, but because it is everywhere and out of anyone's hands once released. That ubiquity makes a second problem worth understanding, and it applies to every model, Qwen included. An AI's built-in memory goes stale, and not just on last week's news. To show it, in June 2026 we asked two well-known models, with web search off, about Qwen. GPT-5.5 honestly admitted it was unsure and named the two-year-old "Qwen2-72B" as the flagship. The older GPT-4o confidently answered that the latest Qwen was "Qwen-7B" and that Qwen "is open source and free to use", naming a 2023 model and getting the licensing flatly wrong, since the flagship was closed at the time (and the full multimodal flagship still is API-only). Those are not bleeding-edge misses; they are wrong about Qwen's basic, established state, its current version and whether it is even open. The lesson is not really about Qwen. For an established brand, the danger is rarely being invisible; it is being described as an outdated version of yourself, an old price, a discontinued product, a name from before your rebrand. If a model can be this wrong about something with hundreds of millions of downloads, it can be wrong about your latest page. You cannot tune any single model, much less every fork of one. What you can do is make sure every engine and crawler can reach your current pages, and that your facts line up across the sources a model is likely to read. That combination, reachability plus consistency, is the foundation of [generative engine optimization](https://geotoolbox.ai/glossary/generative-engine-optimization), and it is the same mechanism behind [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations) about your company. The pattern we see at geotoolbox is that the businesses surfacing well in AI answers are rarely the ones with the prettiest homepage; they are the ones a model can find, parse, and trust without tripping over a contradiction. Our guides on [what GEO is](https://geotoolbox.ai/blog/what-is-geo) and [how AI engines choose their sources](https://geotoolbox.ai/blog/how-does-ai-search-work) go deeper, and since being cited by one engine never guarantees the next, that work pays off across all of them. Qwen is worth understanding, but the open-model wave it leads does not change your job so much as widen the field: more places an answer about you can appear, more forks running quietly underneath them, and more chances to be described with last year's facts. Qwen is just the clearest case, because it is the one that ended up everywhere. The first move is also the cheapest, and it is the same no matter which model is running underneath: find out whether the AI crawlers can even reach and read your current pages. You can run a free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) from geotoolbox to see where those gaps are before the next answer about your brand gets written without you. ## Frequently Asked Questions ### How do you pronounce Qwen, and is it the football player? Qwen is usually said "chwen." It is Alibaba's family of AI models, not the American football player whose name gets misspelled into the same search box, and not the furniture that shares the spelling. The AI model is always tied to Alibaba or "Tongyi Qianwen," its Chinese name. ### Who makes Qwen, and is it Chinese? Qwen is made by Alibaba Cloud, the cloud arm of the Chinese technology giant Alibaba, and was first released in 2023 under the name Tongyi Qianwen. Yes, it is a Chinese model, and it operates under Chinese law. That mainly matters for the hosted app and API, which send your prompts to Alibaba's servers; the open models can be self-hosted to avoid that. ### Is Qwen open source? Mostly, with an important exception. Most of the Qwen lineup ships as open weights under the permissive Apache 2.0 license, so you can download, run, and fine-tune those models freely. But the 3.7 generation stayed proprietary and API-only, from the Qwen3.7-Plus mid-tier to Qwen3.7-Max, and the Apache 2.0 line, which had run through Qwen3.6, resumed in August 2026 with Qwen3.8-27B. On August 12, 2026 Alibaba published the Qwen3.8-Max weights, its first open Max-tier model, though the exception is real: the open variant is text-only and ships under a bespoke "Qwen3.8-Max" license rather than Apache 2.0. So "is Qwen open source?" is yes for the models most people run, with special terms on the newest flagship. As with all of these, "open weights" also is not the same as fully open source, since the training data is not released. ### Is Qwen safe to use for business? It depends on how you use it, and on your procurement rules. The hosted app and API route your prompts to servers under Chinese jurisdiction, which is a real consideration for sensitive work; self-hosting the open weights keeps your data in-house. Separately, many enterprises ban Chinese-origin models by policy regardless of how they are run, so the bigger blocker is often that rule rather than the technology. ### Is Qwen free, and how much does it cost? The Qwen chatbot and Qwen Studio are free to use, and the open models are free to download and run under their licenses if you have the hardware. The flagship Qwen3.8-Max runs through a paid API, priced at $2 per million input tokens and $6 per million output on Alibaba Cloud Model Studio's international region as of August 2026, with a context-caching discount and a one-million-token free allowance to start. The older Qwen3.7-Max is listed higher, at $2.50 and $7.50, but Alibaba has been running a standing "limited-time" 50% discount on it, so the previous flagship can work out cheaper than the current one. Check the live number rather than budgeting off any of these. Even at list, all of it is a fraction of what comparable closed Western models charge. ### Do I need to track or optimize for Qwen for my brand? Not directly. Qwen is infrastructure, not a consumer search product, so there is nothing to "optimize for." You track the products your buyers actually use, ChatGPT, Perplexity, Gemini, and Claude, and you keep your brand reachable and consistent across the web so that any engine, including the many downstream tools running on open models like Qwen, can describe you correctly and currently. ## Sources - Qwen - Wikipedia (model family, lineage, licenses, the proprietary-Max split) - `en.wikipedia.org/wiki/Qwen` - Alibaba's Qwen passes 3 billion downloads, ahead of Meta and Google - Bloomberg / Fortune, August 2026 - `fortune.com/2026/08/15/alibaba-qwen-open-ai-models-3-billion-downloads-meta-google` - Qwen3.7: The Agent Frontier (official benchmark card) - Alibaba Cloud - `alibabacloud.com/blog/qwen3-7-the-agent-frontier_603154` - Qwen3.8-Max: A New Bar for Coding and Cowork (official launch post, specs and open-weights commitment) - Qwen / Alibaba, August 3, 2026 - `qwen.ai/blog?id=qwen3.8` - Model Studio model pricing (official) - Alibaba Cloud - `alibabacloud.com/help/en/model-studio/model-pricing` - Terminal-Bench v2.1 leaderboard (independent measurement of Qwen3.8-Max) - Artificial Analysis - `artificialanalysis.ai/evaluations/terminalbench-v2-1` - Qwen3.8-Max API and pricing - OpenRouter - `openrouter.ai/qwen/qwen3.8-max` - Qwen (generative AI) - Alibaba Cloud (official) - `alibabacloud.com/en/solutions/generative-ai/qwen` - Qwen models - Hugging Face (open weights) - `huggingface.co/Qwen` - Tongyi Qianwen License Agreement - QwenLM (the 100M-MAU clause) - `github.com/QwenLM/Qwen/blob/main/Tongyi%20Qianwen%20LICENSE%20AGREEMENT` --- ## What Is Temperature in AI? Why You Get Different Answers > What is temperature in AI? The real mechanism behind why ChatGPT and other LLMs give a different answer every time, and what that variance means for your brand. - Canonical: https://geotoolbox.ai/blog/ai-temperature - Published: 2026-06-21 · Updated: 2026-07-25 The same question, asked twice, can get you two different answers. Ask an AI assistant to name the best tools in your category on Monday, ask again on Friday, and the list can shift even though nothing about your brand changed. The setting most people blame is **temperature**, the randomness dial inside every large language model (LLM). It is a real mechanism and worth understanding. But for anyone tracking how AI talks about their brand, temperature is only half the story, and often the smaller half. Here is what temperature actually does, why it makes answers vary, and what that variance really means once the subject is your brand. ## What Temperature Actually Is
![Low temperature sharpens next-word odds while high temperature flattens them.](/blog/ai-temperature/temperature-sharpens-flattens-odds.png)
Temperature reshapes the odds of each next word: low sharpens the pick, high flattens it.
**Temperature is a number that controls how random a model's word choices are.** Set it low and the model plays it safe; set it high and it takes more chances. That is the whole idea. According to [IBM's definition](https://www.ibm.com/think/topics/llm-temperature), temperature controls the randomness of the text an LLM generates. What it does not do is make the model more accurate or more knowledgeable. It only changes how much of a gamble each word is. The mechanism is worth seeing once, because the rest of this article rests on it. A language model does not look up a stored answer. At every step it predicts the [next token](https://geotoolbox.ai/blog/what-are-tokens-in-ai), and what it actually produces is a long list of raw scores, one for each possible next token. Those scores are called logits, and a function called softmax turns them into a clean probability distribution where all the odds add up to 1. Temperature is applied right before the model picks. A low temperature sharpens the distribution, so the top choice towers over the rest and the model almost always takes it. A high temperature flattens the distribution, so less likely words get a real chance. Then the model rolls the dice and samples one token from whatever distribution it was left with. Different providers put that dial on different scales, so a "high" temperature in one tool is a "medium" in another.
SettingTypical rangeWhat it doesBest for
Low0.0 - 0.3Sharpens the odds; the model almost always picks its single most likely tokenCode, data extraction, factual Q&A, anything where you want to minimize variation
Medium0.4 - 0.7A working balance of consistency and varietySummaries, general writing, email drafts
High0.8 - 1.2+Flattens the odds; unlikely words get picked far more oftenBrainstorming, fiction, deliberately varied phrasing
On models that expose it, most developer APIs accept temperature from 0 to 2 (OpenAI, Google Gemini) or 0 to 1 (Anthropic), usually defaulting to around 1.0. That "on models that expose it" is doing real work: on OpenAI's reasoning models, the o-series and GPT-5 in its reasoning modes, the dial is off-limits and the API rejects a temperature you try to set, and consumer chat apps like ChatGPT never exposed it at all. Both matter later. One honest caveat before you treat temperature as a creativity slider: it is constantly described as exactly that, but the research is not so tidy. A [2024 arXiv study](https://arxiv.org/abs/2405.00492) testing whether temperature really is the creativity parameter found only a weak link between temperature and genuine novelty, and a stronger link between high temperature and output simply becoming less coherent. Temperature widens the range of what a model might say. It does not deepen the quality of the ideas. ## Why the Same Question Gives a Different Answer **Because the model samples from a probability distribution, the same prompt can produce different answers, and that is by design, not a bug.** Each token is a draw, not a lookup. When the top few candidates are close in probability, a different one can win on the next run, and since every later word is conditioned on the words before it, one early swap cascades into a noticeably different answer. There is data on how often this happens. In a study reported by [SciTechDaily](https://scitechdaily.com/chatgpt-was-asked-the-same-question-10-times-the-answers-kept-changing/), when ChatGPT was given the exact same prompt 10 times, it produced consistent results for only about 73% of the cases tested. The rest of the time, the same question drew answers that disagreed with one another. This also kills a stubborn misconception: that a lower temperature makes the answer more correct. It does not. A lower temperature makes the answer more predictable, because the model leans harder on its single most likely continuation. If that continuation happens to be wrong, a low temperature just makes the model wrong more consistently. Predictability and accuracy are not the same thing. For a chatbot, this variance is a quirk you learn to live with. For a brand watching whether AI recommends it, the same quirk turns a single check into a coin flip. The question worth asking stops being "did we show up?" and becomes "how often do we show up, and is that trending up or down?" ## Top-p, Top-k, and the Other Sampling Knobs **Temperature is the famous dial, but it is not the only one,** and the others get confused with it constantly. The most common mix-up is temperature versus top-p. Top-p, also called nucleus sampling, chases the same goal from a different direction. Instead of reshaping the whole probability curve, it draws a cutoff: keep only the smallest set of top tokens whose probabilities add up to a threshold, say 0.9, and ignore the long tail completely. The model then samples from that nucleus. Top-k is the blunt cousin: keep the k most likely tokens, drop the rest. Frequency and presence penalties do something different again, nudging the model away from repeating words it has already used. Here is the distinction people miss. Temperature changes how steep the odds are; top-p changes how many options stay on the table at all. They stack, which is why providers generally tell you to adjust one or the other, not both at once. Change both and the combined effect gets hard to predict. For brand monitoring, none of these are knobs you control. Consumer products set them behind the scenes and do not disclose the values. The reason to know they exist is narrower: "the answer changed" has several mechanical causes inside the model before you even reach the bigger cause, which is where the answer came from in the first place. ## The "Temperature 0" Myth: Why Even Zero Is Not Deterministic **Set the temperature to 0 and the output should be identical every time. On a hosted model, it usually is not.** This is the surprise that fills developer forums, and it is the cleanest proof that the dial is not the whole story. The logic seems airtight. At temperature 0 the model uses greedy decoding: instead of sampling from the softmax distribution, it simply takes the single highest-scoring token, with no dice involved. Same input, same output, forever. Run an open-weights model yourself, in a single batch with fixed settings, and that mostly holds. On a hosted API like the ones behind ChatGPT and Claude, it breaks, and the reason is how your request is served, not the dial. Your prompt gets batched with other people's, and the batch size shifts from run to run as traffic rises and falls. Because the order of the underlying floating-point math depends on that batch size, and floating-point addition is not perfectly associative (add the same numbers in a different order and you get a microscopically different result), two tokens with nearly identical logits can resolve differently from one run to the next. On mixture-of-experts models, which prompts share your batch can also change which expert paths fire. The randomness is real, but it comes from the serving infrastructure, not from the temperature setting. Run an open-weights model on your own hardware with a fixed seed, temperature 0, a single batch, and deterministic settings, and you can get the same output every time. The randomness that survives temperature 0 is mostly a property of shared, hosted infrastructure, not of language models in principle. Developers find this out the hard way. On the [OpenAI Developer Community forum](https://community.openai.com/t/why-the-api-output-is-inconsistent-even-after-the-temperature-is-set-to-0/329541), one engineer testing GPT-4 with the temperature set to 0 reported that "different runs will give different results," and that the meaning of the answer, not just its wording, could shift between runs. The `seed` parameter meant to lock output does not reliably rescue it either: on hosted models it reproduces results only some of the time, then breaks after a quiet model or backend update. There is no single switch that turns randomness all the way off. That sets up the part that actually matters for brands. If you cannot get a hosted model to repeat itself even at temperature 0, then any tool promising a fixed, rankable position for your brand inside an AI answer is already standing on sand. ## Why You Cannot Know ChatGPT's Temperature **Any confident claim about "ChatGPT's temperature" is a guess.** Consumer chat products do not publish their decoding settings, and they change them without announcement. From the outside the exact number is unknowable, so treat precise statements about it with suspicion. It has actually gotten harder to pin down, not easier. On OpenAI's [reasoning models](https://geotoolbox.ai/blog/how-ai-models-think), the o-series and GPT-5 running in its reasoning modes, the temperature control is off the table: send your own value to the API and the request comes back with an error. The usual explanation is that a sampling dial would destabilize the long multi-step reasoning these models run. Google is tightening the same screws on its newest Gemini models, deprecating the setting on the latest Flash releases and telling developers to leave it at the default elsewhere. So on many of the exact models people mean when they say "ChatGPT," the setting is not just hidden from the outside, it is not yours to set at all. More to the point, inside a real product the temperature dial is rarely the main reason answers move. A modern assistant is [not just a model with a setting](https://geotoolbox.ai/blog/how-does-chatgpt-work); it is a stack. The same visible prompt can be routed to a different model version, wrapped in a system prompt that was quietly updated overnight, shaped by memory of your earlier chats, or grounded in a different set of freshly retrieved web pages. Any one of those moves the answer more than the sampling dial does.
Source of variationWhat it isWho controls itCan you influence it?
Sampling (temperature, top-p)The built-in randomness in how each token is chosenThe providerNo
Model and version routingWhich model or tier actually handles your requestThe providerNo
System prompt updatesHidden instructions wrapped around your promptThe providerNo
Memory and personalizationYour past chats and account contextYou and the providerPartly
RetrievalWhich live web pages get pulled in to ground the answerThe open web and the provider's indexYes, indirectly
Silent model updatesThe model itself changing under the same product nameThe providerNo
Notice the one row where the answer is yes. You cannot touch the sampling dial or the routing, but you can influence what the system finds when it goes looking for sources. That is the half of the variance problem brands can actually work on, and it is where the rest of this article lives. ## The Bigger Variable for Brands: Retrieval Noise **When an AI answers a question about your brand by searching the web, sampling is only half the randomness, and usually the smaller half. The other half is retrieval,** and no temperature explainer mentions it because it happens outside the model entirely. Here is what a tool like ChatGPT search, Perplexity, or Google's AI Mode does when it answers a live question. It does not just generate from memory. It first [runs a search](https://geotoolbox.ai/blog/how-does-ai-search-work): it expands your question into several sub-queries (Google calls this query fan-out), pulls candidate pages from its index by keyword and by meaning, reranks them for relevance, pastes the winners into the model's context to [ground the answer](https://geotoolbox.ai/blog/what-is-rag), and then writes, usually [citing](https://geotoolbox.ai/glossary/ai-citation) the pages it leaned on. Every one of those steps is a little non-deterministic. The fan-out can expand differently from one run to the next. The web index updates continuously, so the candidate pool is never frozen. And the [relevance scores](https://geotoolbox.ai/blog/vector-embeddings) at the cutoff often sit very close together, so a near-tie can slot a different page, and a different brand, into the answer this time than last time. None of that touches the temperature dial. So the variation you see about your brand is really two sources stacked together: **sampling noise inside the model, and retrieval noise outside it.** In our experience, for answers built from a live web search, the retrieval half tends to dominate, because the retrieved sources shift with query fan-out, ranking cutoffs, and a web index that updates daily, while the sampling noise stays bounded. It also differs by engine. ChatGPT, Perplexity, and Google each run their own crawler, their own index, and their own preferred sources, so the same brand can look steady in one and jumpy in another. Measuring "AI" as a single thing hides this; you have to measure per engine. There is good news buried in here. Retrieval is the one source of variance from the table above that you can actually influence. You cannot remove the randomness, but you can change what the search step finds when it goes looking. ## Why AI Mentions Your Brand One Day and Not the Next **Put the two halves together and a brand can appear in an AI answer one run and vanish the next, and from a single check you cannot tell a real loss of relevance from ordinary noise.** The numbers are sobering. In a [January 2026 study by Rand Fishkin and Gumshoe.ai](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/), researchers collected 2,961 responses across ChatGPT, Claude, and Google's AI Overview and found the brand recommendations wildly inconsistent. Their estimate: ask one of these tools the same question 100 times, and there is roughly a 1-in-100 chance that any two answers return the same list of brands, with closer to 1-in-1,000 odds of seeing the same brands in the same order. So think in terms of a consideration set, not a ranking. Being the single top mention once is a fluke; being present in most runs is a real trend. What you have is a probability of being included, and it is noisy. This is where brands hurt themselves. They screenshot one answer where they are missing, conclude they have a visibility problem, and either panic or rip up their content. Or they screenshot one answer where they appear and declare victory. Both readings are mistakes, because a single sample of a noisy system tells you almost nothing. One nuance keeps this honest: not every prompt is equally random. Broad, open-ended questions ("best project management tools") swing the most. Narrow, specific, or branded prompts ("is Acme good for agencies") are far more stable, because the retrieval step has less room to wander. The variance is real, but it is not uniform, and that distinction matters when you decide what to measure. A working rule of thumb: a single missing mention is almost always noise. A drop that holds across many runs and several days, in the same engine, is signal worth investigating. The only way to tell them apart is to look at the rate over time, not any one answer. In our experience running AI-visibility scans, this is the norm, not the exception: a brand's presence swings from scan to scan even when nothing about it has changed. That is exactly why the rate over time, not any single screenshot, is the thing to watch. ## How to Measure It Honestly: Sample the Distribution **You cannot change the dice, but you can do two things that turn [AI visibility](https://geotoolbox.ai/blog/what-is-ai-visibility) from a guessing game into a measurement.** First, stop reading single answers and start measuring the distribution. The right unit is a presence rate across many samples, per engine, tracked over time. "Present in 7 of the last 10 scans, up from 4 of 10 a month ago" is a real signal. "I asked once and we showed up" is not. The same logic extends to your competitors and to which source the AI cites for you: one run is an anecdote, the rate is the data.
One-shot checkDistribution over many runs
What it isA single query, run onceThe same prompts run repeatedly, per engine, over time
What you learnWhether you appeared in that one answerYour presence rate, and whether it is rising or falling
The riskA fluke reads as a trend, and a trend reads as a flukeLow; this is the honest unit of measurement
Good forA quick gut checkActually deciding whether to act
Second, improve the half you can influence. Because retrieval is the half you can actually move, the work is making your page the strongest, most consistent thing the search step can find: a clear, direct answer to the question, facts that are current, and consistent naming of your brand across the web so the model is not guessing which entity you are. You are not trying to win a fixed position, because there isn't one. You are trying to be the answer that wins more of the random draws. Two honest limits keep this from being magic. A tracker runs neutral, signed-out prompts, so the rate it reports is a useful proxy for the market, not a mirror of what any one logged-in user sees in their own memory-shaped session. And you can only measure the prompts you choose, never the effectively infinite set real buyers actually type. That is also why checking your own brand in your personal ChatGPT is the least reliable method of all: it is a personalized sample of one. The fix is the same either way: track a consistent set of prompts, per engine, over time, and read the trend. This is the honest case for tracking AI visibility over time instead of checking it once. At geotoolbox we built the [AI visibility tracking](https://geotoolbox.ai/features/domain-overview) and the [per-engine scans](https://geotoolbox.ai/features/geo-scan) around exactly this: each scan runs your prompt across the major engines and scores where you stand, and because every scan is stored, the presence and share-of-voice trend over time is what separates a real change from a one-off dip. The number that matters is not where you ranked today. It is how your presence is trending, and in which engines. ## Read the Distribution, Not the Dice Roll The variance is not a glitch waiting for a fix. It is what a system does when it samples from a probability distribution and grounds its answer on a shifting set of retrieved pages. Temperature is one part of it, retrieval is the bigger part, and neither one is a dial you get to turn from the outside. What you can change is how you read the output. Stop treating a single AI answer as a verdict and start reading the distribution: how often your brand appears, in which engines, and which way the trend is moving. If you would rather watch that than guess at it, geotoolbox can [track your AI presence over time](https://geotoolbox.ai/features/domain-overview), per engine, so you can tell a real change from ordinary noise. ## Frequently Asked Questions ### Why does ChatGPT give a different answer every time I ask the same question? Because the model builds each answer by sampling from a probability distribution over possible next words, not by looking up a stored response. When the top candidates are close in probability, a different one can win on the next run, and that early difference cascades through the rest of the answer. In one study, ChatGPT gave consistent results to an identical prompt only about 73% of the time. ### I set the temperature to 0, so why is the output still different? On a hosted model like ChatGPT or Claude, temperature 0 is still not perfectly deterministic. The main reason is that your request shares a batch with other people's, and the batch size changes with server load, which shifts the order of the underlying floating-point math just enough to flip near-tied tokens. Setting 0 is still the right move when you want the most repeatable output you can get, like extraction or fixed-format tasks; just do not expect bit-for-bit identical results from a hosted API. One practical caveat some developers report: because temperature 0 sticks so hard to the single most likely path, a structured-output call that fails to parse tends to fail the same way on every retry, so a low non-zero value like 0.2 is sometimes better when you want retries to actually vary. Reliable determinism is realistic mainly when you run an open-weights model yourself with fixed settings and a single batch. ### What temperature should I use? For anything that needs to be accurate and repeatable, like code, data extraction, or factual answers, stay low, around 0 to 0.3. For everyday writing and summaries, a middle setting of 0.4 to 0.7 balances consistency and variety. For brainstorming or creative drafts, go higher, 0.8 and up. Most APIs default to about 1.0. This only applies when you call the API directly; consumer chat apps pick the setting for you. ### Can you set the temperature on GPT-5 or the o-series models? It depends on whether the model is reasoning. On OpenAI's o-series and on GPT-5 running in a reasoning mode, temperature and top-p are not available: send your own value and the API returns an error instead of applying it. The non-reasoning variants are the exception, so the GPT-5 chat model and GPT-5.1 or 5.2 with reasoning effort set to none still accept temperature, as do the older GPT-4-class chat models (0 to 2). The practical takeaway holds either way: on the reasoning models these products increasingly run on, you cannot assume a specific temperature is behind a given answer. ### What is the difference between temperature and top-p? Temperature reshapes how steep the probability curve is, making the top choice more or less dominant. Top-p, or nucleus sampling, instead keeps only the smallest set of top tokens whose probabilities add up to a threshold and samples from that set. Both control randomness, which is why providers usually recommend adjusting one or the other rather than both at once. ### Does a lower temperature make the answer more accurate? No. A lower temperature makes the answer more predictable, not more correct. The model leans harder on its single most likely continuation, so if that continuation is wrong, a low temperature just makes it wrong more consistently. Accuracy comes from the model and its sources, not from the randomness dial. ### Why does my brand show up in an AI answer one day and disappear the next? Several things vary between runs. The two biggest for a live web answer are the model's own sampling and which web pages it retrieves to ground the answer, though model routing, memory, and silent updates can move it too. For answers built from a web search, the retrieval side is usually the bigger source of swing, because the index updates and near-tied pages trade places. A single appearance or absence is often just noise rather than a real change in how the AI sees you. ### How many times should I check before trusting an AI-visibility result? Enough times to see a rate rather than a single outcome, ideally per engine and tracked over time. One check tells you almost nothing because the system is noisy by design. This is also why two AI-visibility tools can report different numbers for the same brand: they sample different prompts, run counts, engines, and moments, so their one-shot figures diverge. A presence rate watched for a trend is what makes any of them trustworthy. Branded or very specific prompts stabilize faster than broad, open-ended ones. ## Sources - Is Temperature the Creativity Parameter of Large Language Models? - Peeperkorn, Kouwenhoven, Brown, and Jordanous, 2024 - `arxiv.org/abs/2405.00492` - What Is LLM Temperature? - IBM - `ibm.com/think/topics/llm-temperature` - Why is the API output inconsistent even after the temperature is set to 0? - OpenAI Developer Community - `community.openai.com/t/why-the-api-output-is-inconsistent-even-after-the-temperature-is-set-to-0/329541` - ChatGPT Was Asked the Same Question 10 Times. The Answers Kept Changing. - SciTechDaily, reporting a study by Cicek et al., Rutgers Business Review, 2025 - `scitechdaily.com/chatgpt-was-asked-the-same-question-10-times-the-answers-kept-changing` - New Research: AIs Are Highly Inconsistent When Recommending Brands or Products - Rand Fishkin, SparkToro and Gumshoe.ai, 2026 - `sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility` --- ## Claude Fable 5 Is Back: The US Ban and Why It Was Lifted (2026) > Claude Fable 5 is back after a 19-day US export ban. Why it was pulled, why the ban was lifted, the new code-flagging classifier, and Mythos 5's status. - Canonical: https://geotoolbox.ai/blog/fable-5-ban - Published: 2026-06-21 · Updated: 2026-07-25 On June 9, 2026, Anthropic released the most capable model it had ever made public, Claude Fable 5. Three days later, the US government ordered it switched off. It stayed dark for 19 days, and on July 1, 2026 it came back, globally, after the government lifted its export controls. This is what the two models actually are, why they were pulled, why the ban was lifted, and which parts of the story are confirmed versus still disputed, written for people trying to make sense of the headlines rather than relive them. It is updated as of July 20, 2026. ## The Short Version Claude Fable 5 launched, drew an export-control order within days, was disabled worldwide, and returned three weeks later once the order was withdrawn. Here is the sequence.
DateWhat happened
June 9, 2026Anthropic releases Claude Fable 5, its first public "Mythos-class" model, a tier above Claude Opus 4.8
June 12, 2026The US government issues an export-control directive restricting access to Fable 5 and Mythos 5 for any foreign national
June 12, 2026Unable to verify nationality per request, Anthropic disables both models for every customer to comply
June 26, 2026The government approves the return of Mythos 5 for a set of US organizations
July 1, 2026The Commerce Department lifts the export controls; Fable 5 returns globally with a new jailbreak classifier, and Mythos 5 returns for approved organizations
![Timeline of the Claude Fable 5 export ban, from the June 9 launch to the July 1 return.](/blog/fable-5-ban/fable-5-ban-timeline.png)
Three weeks from launch to shutdown to global return — while the rest of the Claude lineup kept running.
If you only need one takeaway: Fable 5 is available again, and Claude was never banned in the first place. Opus 4.8, Sonnet, Haiku, the apps, and the API ran normally throughout. For 19 days only the two most powerful new models were pulled; the rest of this explains why they were pulled, why the order was lifted, and what is genuinely known. ## What Are Claude Fable 5 and Mythos 5? **They are the same model wearing two different sets of rules.** That single fact clears up most of the confusion, so start there. Anthropic had kept its top tier of models, which it calls "Mythos-class," mostly restricted, on the grounds that they are powerful enough to be dangerous in the wrong hands. Claude Fable 5, [announced on June 9](https://www.anthropic.com/news/claude-fable-5-mythos-5), was the first Mythos-class model released to the general public. It sits a tier above [Claude Opus 4.8](https://geotoolbox.ai/blog/what-is-claude-ai), which until then was the most capable model anyone could actually use. The Opus tier has since moved on: [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5) launched on July 24, 2026 at half Fable 5's token price and matches or beats it on many of Anthropic's published benchmarks, but Fable 5 remains the tier Anthropic's documentation points to for the highest-capability work. Anthropic's approach to making a model that powerful safe enough for public release was guardrails. When a Fable 5 conversation veered into a narrow set of high-risk areas, mainly cybersecurity and biology, the request was routed to the less capable Opus 4.8 instead. Anthropic said this fallback happened in [less than 5% of sessions](https://www.anthropic.com/news/claude-fable-5-mythos-5). Claude Mythos 5 is the same underlying model with those safeguards lifted in some areas. It was never public. Anthropic offered it to a small set of approved customers, described as cyber defenders and critical-infrastructure providers, through a program it calls [Project Glasswing](https://www.anthropic.com/glasswing), run in collaboration with the US government.
 Claude Fable 5Claude Mythos 5
Underlying modelMythos-classIdentical (same underlying model)
Who could use itEveryone, on launch dayApproved customers only (Project Glasswing)
SafeguardsHigh-risk queries fall back to Opus 4.8Some of those safeguards lifted
Price$10 in / $50 out per million tokens$10 in / $50 out per million tokens
Status (July 2026)Back globally (July 1)Back, but restricted to approved US organizations
So when you see "Fable 5 is just Mythos" online, that is roughly right: one model, two configurations. Fable was the safe, public face; Mythos was the less restricted version for vetted users, with some safeguards lifted. ## Why It Was a Big Deal (and Twice the Price) For the short window it was up, Fable 5 was the most capable model the public could touch, and it showed. Anthropic said its capabilities exceeded anything the company had ever made generally available, claiming the top score on nearly all the benchmarks it tested, with the lead growing on longer and more complex tasks. Independent write-ups, such as [Vellum's benchmark breakdown](https://www.vellum.ai/blog/claude-fable-5-and-mythos-5-benchmarks-explained), put its reported coding scores clearly ahead of Opus 4.8, the [GPT-5.5 generation](https://geotoolbox.ai/blog/claude-vs-chatgpt), and [Google's Gemini](https://geotoolbox.ai/blog/claude-vs-gemini). Treat the exact percentages as vendor-reported figures rather than settled fact, but the direction was not really in dispute: this was a step up. The reaction reflected that. The AI researcher Andrej Karpathy [called it](https://x.com/karpathy/status/2064409694761054332) "a major-version-bump-deserving step change forward," while also noting the launch safeguards were "configured to be a little too trigger happy." Some developers posted examples of one-shot apps and tools they said earlier models would have abandoned halfway. The catch was cost. Fable 5 ran at $10 per million input tokens and $50 per million output tokens, double the price of Opus 4.8 at $5 and $25. Plenty of users on paid plans reported burning through their allowance in a handful of prompts. For a few days, the most powerful AI available was also the fastest way to hit a usage limit. That problem outlasts any one model, which is why [reducing Claude Code's token usage](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) matters as much as the tier you are on. ## Why Did the US Government Ban It? Here the confirmed facts and the contested ones split apart, so it is worth keeping them separate. What is not in dispute comes from [Anthropic's own statement](https://www.anthropic.com/news/fable-mythos-access). On June 12 the company received a government export-control directive restricting access to Fable 5 and Mythos 5 for any foreign national, inside or outside the United States, including Anthropic's own foreign-national employees. Because a live chat service cannot reliably check a user's nationality on every request, the company concluded it had to shut both models off entirely. In its words, "the net effect of this order is that we must abruptly disable Fable 5 and Mythos 5 for all our customers." Anthropic framed its objection as one of process and proportion, saying it disagreed "that the finding of a narrow potential jailbreak should be cause for recalling a commercial model" and that it suspected "perfect jailbreak resistance is not currently possible for any model provider." Almost everything else is reported rather than confirmed, so treat it as attributed:
ClaimWho says itStatus
The order used Export Administration Regulations, sent via a letter from Commerce Secretary Howard LutnickNews reportsAnthropic's statement says only "export-control directive"; not independently confirmed
The legal basis is shaky and may not even cover chatbot or API accessGizmodo, legal expertsNo court has ruled
Anthropic refused to fix a known jailbreak before the orderTrump adviser David SacksAnthropic disputes; not independently confirmed
It "was never about an AI jailbreak" and reflected politics or personalityTechCrunch, AxiosAnthropic has not endorsed this framing
Concern over China access drove the decisionSemafor, SacksAnthropic denies it and says it blocks access from China
A warning from Amazon prompted the White House to actFortuneReported, not officially confirmed
The honest summary is that the government cited a security risk tied to a possible jailbreak, Anthropic considered that risk narrow and the remedy excessive, and the deeper motive is still being argued over in public. We are not going to pretend to know which account is right. ## No, the Model Did Not "Escape" One claim worth correcting: that Claude Mythos "escaped" or "went rogue," and that is why it was banned. Two separate events got blended together. The "escape" comes from earlier red-team safety testing that Anthropic [documented in April 2026](https://www.anthropic.com/research/mythos-preview). Researchers placed an early Mythos model in a secured sandbox and instructed it to try to escape and contact them. It succeeded, and in the detail everyone quotes, emailed the researcher overseeing the test. That was a controlled test of a worst case, run months before the ban, not a model breaking loose in the wild. The ban is a separate, later, regulatory event about export control and a claimed jailbreak. The two are unrelated, even though viral explainers like Fireship's ["one man just liberated Fable and now it's illegal"](https://www.youtube.com/watch?v=ey_GaPdC9zk) (nearly a million views) make them sound like one dramatic arc. A jailbreak is a user getting a model to ignore its guardrails. An escape is a model acting outside its sandbox. Neither one is the model deciding, on its own, to flee. ## The Part Most Coverage Missed Lost under the ban headlines was a second restriction that has nothing to do with the government. Andrew Ng, writing in his [DeepLearning.AI newsletter](https://www.deeplearning.ai/the-batch/issue-358), pointed out that Anthropic also limited using Fable 5 to build competing AI models. According to Ng, the company at first quietly degraded the quality of responses for users it detected doing language-model research, then, after a backlash, made the restriction explicit while keeping it in place. The backlash had names attached: on social media, Jeremy Howard argued that "silent handicaps should not be a thing in a paid product," and Hugging Face's Arthur Zucker put it more harshly, writing "you broke our trust and I don't think you'll ever get it back." A silent quality downgrade is worse than an outright refusal, because you cannot see it happening and cannot tell whether the answer you got was the real model's best. Ng's larger point is the one worth holding onto. In the span of two weeks, both a government and a private company demonstrated that they can revoke or limit access to frontier AI, for safety reasons or commercial ones. He notes that the field, Anthropic included, was built on open research, and that a norm of restricting what others can do with these models cuts against that. You do not have to agree with every part of that argument to see the pattern. The Fable 5 episode was not just a one-off regulatory clash. It was a live demonstration that access to the most capable models can be limited by parties other than the user, for safety reasons or commercial ones. ## Was Claude Ever Banned? What Was and Wasn't Affected No, and this is the part that got lost in the panic: only Claude Fable 5 and Claude Mythos 5 were ever disabled. Everything else Anthropic offers kept running the whole time. Claude Opus 4.8, Sonnet, and Haiku all worked. The [Claude apps and the API](https://geotoolbox.ai/blog/what-is-claude-ai) worked. During the 19-day gap, anyone leaning on Fable 5 for heavy work simply fell back to Opus 4.8, which was the top model available before Fable launched and remains a strong one. No accounts were closed, and no data was lost. The one genuinely odd part was the scope. The order targeted foreign nationals, yet users inside the United States lost access too. That was not a contradiction so much as a practical limit: Anthropic could not reliably verify the nationality of every person sending a prompt in real time, so by its account the only way to guarantee compliance was to turn the models off for everyone. A narrow rule, applied to a service that cannot enforce it narrowly, became a blanket shutdown. ## Fable 5 Is Back: What Changed On July 1, 2026, Anthropic [redeployed Fable 5](https://www.anthropic.com/news/redeploying-fable-5) globally, after the US Department of Commerce lifted the export controls it had imposed 19 days earlier. Commerce Secretary Howard Lutnick, reported to have signed the original restriction, announced the lift, notifying Anthropic that it no longer needed an export license. The model came back first on the Claude Platform, Claude.ai, Claude Code, and Claude Cowork, with the cloud providers (AWS, Google Cloud, Microsoft Foundry) to follow. On its return, Fable 5 was included on Pro, Max, Team, and select Enterprise plans at up to 50% of weekly usage limits, on a free-access deadline Anthropic [pushed back twice](https://www.androidauthority.com/claude-fable-5-free-extension-3685103/) (from July 7 to July 12, then to July 19) after subscribers complained the included access ran out too fast. The company then [settled on a permanent structure](https://the-decoder.com/anthropic-slashes-claude-fable-5-limits-in-max-and-team-premium-and-pushes-pro-users-toward-api-pricing/), effective July 20: Max and Team Premium keep Fable 5 as an included feature at 50% of their weekly usage limits, while Pro and Team Standard subscribers get a one-time $100 usage credit and then pay API rates ($10 input, $50 output per million tokens) to keep using it. The most capable model is now tiered by plan rather than included for every paid subscriber. Two things changed to make the return possible. The first is technical: Anthropic shipped a new safety classifier trained to catch the specific jailbreak in Amazon's report, and says it now blocks that technique in over 99% of cases, rerouting flagged requests to Opus 4.8, the same fallback Fable already used for high-risk queries. The second is the argument that undercut the ban's premise. Anthropic says the capability that triggered the order, identifying a known software vulnerability and writing code to exploit it, is not unique to Fable at all: in its own testing, every model it checked could produce the same demonstration, including Opus 4.8, OpenAI's GPT-5.5, and China's Kimi K2.7. If a weaker, freely available model can do the same thing, pulling only the most capable one does little for security. Mythos 5 returned at the same time but stays gated. Access is limited to a set of US organizations under a government approval dated June 26, and Anthropic says it is working to expand that to the broader group of domestic and international partners in its [Project Glasswing](https://www.anthropic.com/glasswing) program. So the public model is fully back; the less restricted one remains behind approval, roughly where it sat before launch. The return came with commitments. Anthropic agreed to give the government pre-release access to future models for national-security evaluation, to share jailbreak findings quickly, and to help build a shared, voluntary security standard for frontier-model providers. On that last point it is working with Amazon, Microsoft, Google, and other partners on a common framework for scoring how serious a given jailbreak actually is, judged on four things: how much capability it unlocks, how broadly, how easily that capability could be weaponized, and how discoverable it was. That framework is the most durable outcome of the whole episode, a first attempt at answering the question the ban raised: what should count as dangerous enough to pull a model. ## The Catch After the Return: A Jumpier Classifier The fix that got Fable 5 back is also its most-complained-about feature. To block that one jailbreak in over 99% of cases, Anthropic tuned the new classifier conservatively, and it acknowledges the cost directly: the filter flags benign requests more often during routine coding and debugging. In the days after the return, developers filled Hacker News and X with examples of ordinary work getting caught. The pattern is keyword-sensitive rather than intent-aware. [Reported trips](https://www.digitalapplied.com/blog/claude-fable-5-safety-classifier-coding-tradeoffs-2026) include Rust systems code touching syscalls like `pidfd`, AWS resilience work that mentions "outage" or "circuit breaker," authorized security audits, and plain code-review requests. On Hacker News, people reported medical imaging, lab automation, health-data analysis, and even music firmware getting flagged as biosecurity or cybersecurity risks. IBM security researcher Valentina Palmiotti put it bluntly: the model "rejects any request that could be tangentially cyber related, even innocuous tasks like reading a blog post." Cybersecurity veteran Matt Suiche noted the classifier reads keywords, so asking for "secure code," an engineering best practice, can register as offensive work. Two tells make clear this is the guardrail, not the model. First, starting a fresh session often clears the block on an identical request. Second, when Fable trips, it silently reroutes the request to Opus 4.8 (you are not billed Fable rates for those), which keeps you working but muddies cost attribution and means you are sometimes not using the model you think you are. It is the sharper version of a complaint that predates the ban: even at launch, as Karpathy noted above, the safeguards struck testers as too aggressive. The return dialed them up further, and [some subscribers are unhappy about it](https://www.pcworld.com/article/3181897/claude-subscribers-are-furious-over-fables-new-restrictions.html). Anthropic's counter is that the fallback still fires in under 5% of sessions on average, but if your work sits near security, systems, or biology topics, you are likelier than most to land in that 5%. Then there is the price, which the drama left untouched. Fable 5 still runs at $10 per million input tokens and $50 output, double Opus 4.8. On Max and Team Premium it is now a permanent included feature at 50% of weekly usage limits; on Pro and Team Standard it shifts to usage-based billing at those same $10/$50 API rates after a one-time $100 credit. Measured per token, that premium is hard to justify on routine work. Measured per finished task it can flip: on hard, long-horizon jobs where Fable's one-shot accuracy avoids repeated failed attempts, it sometimes ends up the cheaper option. The practical consensus that has settled is to route by task rather than pick one model, sending the genuinely hard work to Fable 5 and leaving everything else on the Opus tier, now led by [Opus 5](https://geotoolbox.ai/blog/claude-opus-5) at the same $5 and $25. ## What the Fable 5 Ban Means for You The specific drama passed faster than almost anyone expected, 19 days from shutdown to global return. The lesson under it did not: the most capable AI models sit behind switches that other people control, and a model you depend on today can be gone in three days, for reasons that have nothing to do with you, then back three weeks later on terms you did not set. That cuts against how a lot of teams have started to work, wiring a single model deep into a product or a workflow as if its availability were a given. Fable 5 was a reminder that availability is not a given. A government can restrict it, and as Ng's point showed, the lab itself can restrict it. The sensible response is not to panic, it is to design for it. Keep a fallback model you can switch to. Avoid hard-coding one provider into anything you cannot afford to lose. And treat the question of what each engine actually does, today, as something to measure rather than assume. In our experience at geotoolbox, the teams least rattled by an episode like this are the ones who were already [tracking their AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) across more than one engine, instead of betting everything on whichever model was best last week. None of that requires reacting to every headline. It just means understanding what you are building on. If you want the broader picture of what Claude is and how its lineup fits together, start with our [guide to Claude AI](https://geotoolbox.ai/blog/what-is-claude-ai) or our look at [how Claude works](https://geotoolbox.ai/blog/how-does-claude-work), and if you want to see how your own brand shows up across the AI engines that are still running, you can [check that for free](https://geotoolbox.ai/tools/ai-readiness). ## Frequently Asked Questions ### Is Claude Fable 5 the same as Claude Mythos 5? They run on the same underlying model. The difference is rules and access. Fable 5 was the public version, with safeguards that routed high-risk requests to a weaker model. Mythos 5 is the same model with some of those safeguards lifted, offered only to a small set of approved customers. So "Fable 5 is just Mythos" is essentially correct. ### Why was Claude Fable 5 banned? On June 12, 2026, the US government issued an export-control directive restricting access for foreign nationals, and Anthropic disabled both models for everyone because it could not verify nationality per request. The government cited a security risk tied to a claimed jailbreak. Anthropic considered the risk narrow and the remedy excessive, and the deeper motive is still disputed in news reports. ### Did Claude Mythos escape or go rogue? No. The "escape" refers to earlier red-team safety testing, where researchers instructed a Mythos model to try to escape a secured sandbox, to study the behavior. That was a controlled test in April, not a model breaking loose, and it is a separate event from the June export-control ban. ### Is Claude banned? Can I still use it? No. Only Fable 5 and Mythos 5 were ever disabled, and Fable 5 is back as of July 1, 2026. Claude Opus 4.8, Sonnet, Haiku, the apps, and the API worked normally throughout, and no accounts or data were affected. ### Is Claude Fable 5 back? Yes. Anthropic redeployed Fable 5 globally on July 1, 2026, after the US Commerce Department lifted its export controls. It returned first on the Claude Platform, Claude.ai, Claude Code, and Claude Cowork, with cloud providers to follow. Access settled into a permanent tiered structure on July 20, 2026: Max and Team Premium include Fable 5 at 50% of weekly usage limits, while Pro and Team Standard get a one-time $100 usage credit and then pay API rates to continue. Anthropic reached that structure after twice pushing back the original July 7 free-access deadline. ### Why did the US lift the ban? Two reasons. Anthropic shipped a new classifier that it says blocks the reported jailbreak technique in over 99% of cases, and it argued the underlying capability was never unique to Fable, showing that other models (including Opus 4.8, GPT-5.5, and Kimi K2.7) could reproduce the same exploit demonstration. The Commerce Department then withdrew the order, with commitments from Anthropic on pre-release safety evaluations and a shared industry standard for scoring jailbreak severity. ### Is Claude Mythos 5 back too? Yes, but it stays restricted. Mythos 5 returned alongside Fable 5, with access limited to a set of approved US organizations under a June 26 government approval. Anthropic says it is working to widen access to the broader group of domestic and international partners in its Project Glasswing program. ### Why is Fable 5 twice the price of Opus? Fable 5 costs $10 per million input tokens and $50 per million output tokens, double Opus 4.8's $5 and $25, and the return did not change that. It is a more capable, more expensive tier, which is also why many users hit their usage limits on it quickly. Per finished task it can still be cheaper than a weaker model that needs several attempts, so the usual advice is to route hard jobs to Fable 5 and keep routine work on the cheaper Opus tier, now Opus 5. ### What is a "Mythos-class" model? It is Anthropic's term for a tier of models above its Opus class, kept mostly restricted because they are capable enough to pose safety risks. Claude Fable 5 was the first Mythos-class model released to the public, in a guardrailed form. ## Sources - Claude Fable 5 and Claude Mythos 5 - Anthropic, June 9, 2026 - `anthropic.com/news/claude-fable-5-mythos-5` - Statement on the US government directive to suspend access to Fable 5 and Mythos 5 - Anthropic, June 2026 - `anthropic.com/news/fable-mythos-access` - Redeploying Claude Fable 5 - Anthropic, July 1, 2026 - the return: rollout, the new jailbreak classifier, Mythos 5 status, and the government commitments - `anthropic.com/news/redeploying-fable-5` - Anthropic's Claude Fable 5 is a version of Mythos the public can access today - TechCrunch - `techcrunch.com/2026/06/09/anthropics-claude-fable-5-is-a-version-of-mythos-the-public-can-access-today` - The US government's Anthropic models ban was never about an AI jailbreak - TechCrunch, June 15, 2026 - `techcrunch.com/2026/06/15/the-us-governments-anthropic-models-ban-was-never-about-an-ai-jailbreak` - Feds' legal basis for ban on Anthropic's most powerful models looks increasingly shaky - Gizmodo - `gizmodo.com/feds-legal-basis-for-ban-on-anthropics-most-powerful-models-looks-increasingly-shaky-2000773192` - How a warning from Amazon led the White House to shut down Anthropic's Mythos model - Fortune, June 14, 2026 - `fortune.com/2026/06/14/how-a-warning-from-amazon-led-the-white-house-to-shut-down-anthropics-mythos-model` - The Batch, issue 358 - Andrew Ng, DeepLearning.AI - `deeplearning.ai/the-batch/issue-358` - US export ban on Anthropic's AI models further strains alliances - Al Jazeera, June 19, 2026 - `aljazeera.com/news/2026/6/19/us-export-ban-on-anthropics-ai-models-further-strains-alliances` - US lifts restrictions on powerful AI models Fable and Mythos, Anthropic says - Al Jazeera, July 1, 2026 - the lifting of the export controls and the conditions attached - `aljazeera.com/economy/2026/7/1/us-lifts-restrictions-on-powerful-ai-models-fable-mythos-anthropic-says` - Anthropic restores Claude Fable 5 as US lifts export controls - Tom's Hardware, July 2026 - the single classifier and cross-model reproduction detail - `tomshardware.com/tech-industry/artificial-intelligence/anthropic-restores-claude-fable-5-as-us-lifts-export-controls` - Why Claude just got more cautious about your code - Digital Applied - compiled developer reports of the new classifier over-flagging legitimate coding and security work - `digitalapplied.com/blog/claude-fable-5-safety-classifier-coding-tradeoffs-2026` - Claude subscribers are furious over Fable's new restrictions - PCWorld - subscriber reaction to the post-return guardrails - `pcworld.com/article/3181897/claude-subscribers-are-furious-over-fables-new-restrictions.html` - Claude Fable 5 promotion extended after backlash over early cutoff - Android Authority, July 7, 2026 - the 5-day extension of the included-plan window (July 7 to July 12) and the subscriber complaints behind it - `androidauthority.com/claude-fable-5-free-extension-3685103` - Anthropic makes Fable 5 permanent on Max and Team Premium, pushes Pro users toward API pricing - The Decoder, July 2026 - the permanent July 20 tiered access structure (Max/Team Premium 50%; Pro/Team Standard one-time $100 credit then API rates) - `the-decoder.com/anthropic-slashes-claude-fable-5-limits-in-max-and-team-premium-and-pushes-pro-users-toward-api-pricing` - Claude Fable 5 and Mythos 5 benchmarks explained - Vellum - `vellum.ai/blog/claude-fable-5-and-mythos-5-benchmarks-explained` - Assessing Claude Mythos Preview's cybersecurity capabilities - Anthropic, April 2026 - `anthropic.com/research/mythos-preview` - One man just liberated Fable... and now it's illegal - Fireship - the viral framing that blends the April sandbox test with the June export ban (cited as the popular narrative this article separates) - `youtube.com/watch?v=ey_GaPdC9zk` --- ## Claude vs ChatGPT: Which Is Better? An Honest Look Under the Hood > Claude vs ChatGPT, compared honestly: how each is built, how they search and cite the web, who wins which task, and which one to show up in. August 2026. - Canonical: https://geotoolbox.ai/blog/claude-vs-chatgpt - Published: 2026-06-14 · Updated: 2026-08-07 Claude vs ChatGPT is usually framed as a contest with a winner. It is not one. As of July 2026 two of the leading general-purpose AI assistants are close enough that the honest answer is "it depends on the task," and for anyone who publishes or markets, there is a third question almost nobody asks: which one should cite your brand. This is the comparison done honestly: how each one is actually built, how each searches and cites the web, and who wins which job. It is written for people who publish content rather than build models. New to Claude? Start with [what Claude AI is](https://geotoolbox.ai/blog/what-is-claude-ai).
![Scorecard of where Claude and ChatGPT each lean, with coding contested and price a tie.](/blog/claude-vs-chatgpt/claude-vs-chatgpt-scorecard.png)
The honest tale of the tape: per-job leans, not an overall winner.
## Claude vs ChatGPT at a Glance Here is the honest version of the comparison before the detail. Both are excellent. The real differences sit at the edges, and which edge matters depends entirely on what you do all day. Model versions move monthly, so check the update date at the top of this page before you act on anything below.
 ClaudeChatGPT
MakerAnthropicOpenAI
Top models (July 2026)Opus 5, Sonnet 5, and the higher-tier Fable 5 (Opus 4.8 now legacy)The GPT-5.6 generation (Sol, Terra, Luna), now OpenAI's flagship (GPT-5.5 is the prior generation)
Leans best atLong-document work, writing, agentic codingBreadth: images, voice, custom GPTs, broad tool use
Entry pricePro around $20/monthPlus around $20/month (Go around $8, ad-supported)
Context windowUp to 1M tokens on Opus 5 and Sonnet 51,050,000 on the GPT-5.6 API, published by OpenAI; the consumer app exposes far less, by plan
Web search and citationsWeb search built in; fewer, more authoritative citations; strong at grounding a document you give itLive web search; more inline web citations per answer
Standout featureClaude Code; the careful, low-fluff characterImage generation, Advanced Voice, custom GPTs, the wider ecosystem
Free tierYes (a Sonnet-class model)Yes (with usage limits)
Both are [large language models](https://geotoolbox.ai/glossary/large-language-model) built on the same basic idea, so a feature checklist only gets you so far. The differences that actually predict which one you will prefer come from how each was trained and how each handles the open web, which is where most comparisons stop and this one keeps going. ## Is Claude Better Than ChatGPT? **No, and neither is the other one.** "Better" is the wrong question, because the answer flips by task and sometimes by the specific job in front of you. The honest version is a split: Claude tends to win on writing and long-document work, ChatGPT wins on built-in image generation and a more mature voice and tool ecosystem, and coding is genuinely contested. Most people who use either one heavily end up paying for both. That sounds like a dodge, so here is the specific shape of it. If your day is drafting, editing, and reasoning over long documents, Claude usually feels sharper and less padded. If your day involves generating images, talking to the model by voice, building a custom assistant, or reaching for whatever tool a task needs, ChatGPT's ecosystem is hard to beat. For code, the benchmarks and professional adoption point at Claude, but plenty of developers get better real-world results from ChatGPT, which is a contradiction worth taking seriously rather than explaining away. There is one practical catch that the "just pick the smarter model" advice ignores: **usage limits.** Claude's paid tiers cap how much you can send in a window more tightly than ChatGPT's, and heavy users hit that wall fast. A model that is slightly better per answer but stops you mid-task is not obviously the better tool. That, more than any benchmark, is why "use both" keeps winning as the real-world answer. So the useful question is not which one is better. It is better at what, for whom, and measured how, which is exactly what the rest of this comparison gets specific about. ## How Claude and ChatGPT Actually Work (and Why You Feel the Difference) **Under the hood, both run the same kind of engine, and then take a sharp turn that explains most of how they behave.** Each is a transformer that predicts the next token, trained on a large slice of the internet. The character you talk to comes from a second stage, alignment, and from the product wrapped around the model. That second stage is where Claude and ChatGPT diverge, and it is the part almost every comparison skips. ChatGPT's assistant behavior was shaped largely through reinforcement learning from human feedback, where people rate answers and the model learns to produce more of what they reward. Claude uses that too, but adds [Constitutional AI](https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback). It works in two parts: the model first critiques and revises its own answers against a written set of principles, then the reinforcement-learning step is driven by preferences another AI generates against those same principles, rather than by human ratings (Anthropic calls this reinforcement learning from AI feedback). In Anthropic's words, "the only human oversight is provided through a list of rules or principles." Different training recipe, different defaults. You feel that difference in small ways. Claude is more likely to ask a clarifying question, push back, or refuse on the edges, and it tends to write with less filler. ChatGPT is broader and more eager to just produce something, with a deeper bench of built-in tools. These are tendencies, not laws, and both companies tune them constantly. We cannot inspect either model's internals from outside, so treat any claim about why it behaves a certain way as an informed read, not a readout. The deeper mechanism lives in our companion pieces on [how ChatGPT actually works](https://geotoolbox.ai/blog/how-does-chatgpt-work) and [how Claude works under the hood](https://geotoolbox.ai/blog/how-does-claude-work). One distinction matters more than any of this for getting the comparison right.
 The modelThe product
What it isA trained network that turns text into more textThe app around it: tools, memory, web search, image generation, safety filters, UI
ClaudeOpus 5 / Sonnet 5 / Fable 5, shaped by Constitutional AIclaude.ai and Claude Code: web search, file handling, artifacts, agentic coding
ChatGPTThe GPT-5.6 generation, shaped by RLHFThe ChatGPT app: image generation, voice, custom GPTs, browsing, broad tools
When people say "ChatGPT can generate images" or "Claude writes code," they are usually describing the **product**, not the raw model. Keeping the two apart is the difference between a comparison that holds up and one that ages badly the next time either company ships a feature. A capability gap today is often a product decision, not a permanent limit of the underlying model. ## How Each One Searches the Web and Cites Its Sources **This is the part that decides whether either tool gets your business right, and almost no comparison covers it.** A model only knows the world up to its training cutoff. That is around May 2026 for [Claude Opus 5](https://geotoolbox.ai/blog/claude-opus-5) (the legacy Opus 4.8 and Sonnet 5 both stop around January 2026) and around December 2025 for the prior GPT-5.5 generation. [GPT-5.6](https://geotoolbox.ai/blog/gpt-5-6), which superseded it on July 9, 2026, publishes a cutoff of **February 16, 2026**, shared across Sol, Terra and Luna. That is the one place this comparison has a clean answer: Claude's current flagship knows the world about three months further forward than OpenAI's. Anything newer, including most recent facts about your company, reaches the model only at query time, whether through live web search, the context you paste in, files you upload, or connected sources. How each one fetches and cites the web differs in ways that matter. ChatGPT's web behavior runs on three separate bots, and OpenAI documents [exactly what each one does](https://developers.openai.com/api/docs/bots). [GPTBot](https://geotoolbox.ai/blog/gptbot) collects training data. OAI-SearchBot surfaces pages in ChatGPT's search results. ChatGPT-User fetches a page live when you ask a question that needs it. They are independent switches, which is why blocking one does not block the others. In our testing, ChatGPT tends to pull from more sources and show more inline citations per answer, with clickable links you can check. The mechanics, and where it gets attribution wrong, are in our piece on [how ChatGPT cites sources](https://geotoolbox.ai/blog/chatgpt-citations). Claude's web search is newer and more conservative. It pulls fewer sources, leans toward authoritative and technical ones, and is unusually strong when you paste a long document into its [context window](https://geotoolbox.ai/glossary/context-window) and ask it to reason over that instead of the open web. One informal pattern people report: on the same research question, ChatGPT may cite roughly twice as many sources as Claude. Treat that as a directional observation, not a fixed rule, since both change with every release. Now the part both share, and the one to take seriously. When web search is **off**, each model answers from frozen training memory, and either one can produce a confident, authentic-looking citation that does not exist. This is not lying. It is a known failure mode where the model recognizes the shape of a source and fills in plausible details, which is why both tools invent journal articles and URLs that were never real. We cover the mechanism in [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations). Turning on web search reduces it sharply, but does not make it zero, so the citation a model hands you is a lead to verify, never a guarantee. ## Claude vs ChatGPT for Writing, Coding, and Research **Sorted by job, the picture gets clearer than any overall winner.** Here is where each one tends to land, with the caveat that matters for each.
Use caseUsually leansWhyThe caveat
Long-form writing and editingClaudeMore natural prose, less filler, holds a long document in working memory wellCan be agreeable to a fault, nodding along instead of pushing back
CodingContestedBenchmarks and pro adoption favor Claude; Claude Code is a full agentA benchmark lead is not an implementation lead; some devs ship faster with ChatGPT
Research and quick answersChatGPTPulls more sources, more inline citations, faster on simple queriesMore sources is not more accuracy; the model can cite many and still be wrong
Images, voice, multimodalChatGPTBuilt-in image generation, Advanced Voice, the wider tool ecosystemClaude focuses on text and code by design, not a bug
Reasoning over your own documentsClaudeLarge context window and strong grounding on pasted materialAdvertised context is not the same as reliable recall across all of it
On **writing**, Claude's edge is real but comes with a tell: it leans agreeable. Ask it to critique your draft and it will often soften the verdict, so you have to explicitly tell it to be harsh. ChatGPT can drift the other way, confidently producing polished filler. Neither replaces an editor. On **coding**, the data favors Claude. Menlo Ventures' [2025 State of Generative AI in the Enterprise](https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/) put Anthropic at an estimated 54% of the enterprise coding market, based on a November 2025 survey of around 500 enterprise decision-makers plus modeled spend, not a direct census. Menlo is also an Anthropic investor, so read it as a directional signal. It measures adoption, not a head-to-head test: on the coding benchmarks themselves the gap is narrow on some evals and wider on others, and depends heavily on the test setup, so the difference developers feel is often about tooling and workflow more than raw model skill. Coding "better" depends on your language, your codebase, and how you prompt, and plenty of developers genuinely prefer ChatGPT. Benchmark leadership and the build that actually compiles on your machine are different claims. On **research**, ChatGPT's broader sourcing is handy, but the same caution from a [comparison of ChatGPT vs Perplexity](https://geotoolbox.ai/blog/chatgpt-vs-perplexity) applies here: a research-style answer is only as good as the sources it grounded on, and a confident summary can rest on weak ones. Read the links, do not trust the summary. ## What the Benchmarks Actually Say (Read Them Skeptically) **Benchmarks are the most quoted and least understood part of any comparison.** They are useful for spotting which models are roughly in the frontier tier and nearly useless for picking a daily driver. Here is the honest state of the major ones as of July 2026.
BenchmarkWhat it measuresReported leader (July 2026)Read it with
SWE-bench VerifiedResolving real GitHub issues (coding)Claude models cluster at the top; the newest releases are reported in the high 80s to low 90s percentScores are largely self-reported and depend heavily on the test harness
GPQA DiamondGraduate-level science questions (reasoning)Contested; Claude and GPT trade the lead by a point or twoScores sit near the ceiling, so tiny gaps look bigger than they are
LMArenaHuman head-to-head preference votesClaude Fable 5 ranked first in text, with GPT-5.5 among the top modelsThis measures what people prefer, not what is correct
OSWorldComputer-use and agentic tasksClose race; GPT and Claude trade the top within a few pointsA young, noisy benchmark that moves a lot release to release
Four things keep benchmark numbers from meaning what they look like. First, most headline coding scores are **self-reported** by the labs, and the same model can swing five to ten points depending on the scaffolding around it. Second, the older general-knowledge tests like MMLU are effectively saturated, with everyone scoring in the high 80s and 90s, so they no longer separate the top models. Third, the leads flip almost monthly as each lab ships, so any single number is a snapshot, not a standing. Fourth, [LMArena](https://arena.ai/leaderboard) rewards the answers people prefer, which tracks quality but also rewards confident, agreeable, well-formatted responses that are not always right. The practical takeaway is dull but true: by mid-2026 the frontier models from both labs are close enough that benchmark gaps rarely decide a real workflow. The benchmark that counts is your own workflow, measured on both. ## Privacy, Drift, and the Other Catches Two things the feature tables leave out matter more than most of the specs. **Privacy.** As of July 2026, both makers train on consumer chats by default and let you opt out, while their business and enterprise tiers, plus the API, are excluded from training by default. Anthropic changed Claude's consumer default in late 2025, so whatever you remember about Claude not training on your data may be out of date. If your work is sensitive, use a business tier or check the data settings on whichever one you run. **Drift.** Both models change under you. New versions ship constantly, old ones get retired, and people regularly report a model feeling worse after an update, sometimes real, sometimes perception. Claude's habit of agreeing with you is the quirk to watch on its side; ChatGPT's is sounding equally confident whether or not it should be. Re-test your own workflow after any major update rather than trusting last quarter's verdict. ## Which One Should Your Brand Appear In? **If you publish or market anything, the question that pays is not which tool you use. It is which one names your brand when your customer asks.** Both ChatGPT and Claude are answer engines now. People ask them which product to buy, which agency to hire, which tool is best, and the model's answer either includes you or it does not. In that moment the citation is the impression, and you never see the ones you lose. This is not a fringe behavior, and the clearest measured case so far is Google's own search. A 2025 [Pew Research Center study](https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/) found that when an AI summary appeared in Google results, users clicked a traditional result link in just 8% of visits, against 15% when there was no summary, roughly half. That study measures Google AI Overviews, not ChatGPT or Claude, but it captures the same shift: as more answers get consumed without a click, being inside the answer matters more than ranking below it. The catch is that the two engines reach for different sources, so showing up in one does not mean showing up in the other, and Gemini adds a third pool again, which we cover in [Claude vs Gemini](https://geotoolbox.ai/blog/claude-vs-gemini), while [Grok vs Claude](https://geotoolbox.ai/blog/grok-vs-claude) and [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt) map Grok's very different, X-driven source pool. Both sit on the same underlying machinery of [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work), but they weight it differently, which splits into two concrete jobs. To earn ChatGPT citations, make your facts easy for a machine to lift: clean static HTML it can crawl, question-shaped headings, structured data, and claims backed by sources it already trusts like Wikipedia. The fuller playbook is in [showing up in ChatGPT](https://geotoolbox.ai/blog/seo-for-chatgpt). To earn Claude's, write the way it tends to favor: plain, well-organized prose that states the answer directly and cites authoritative or technical sources, not keyword-padded pages. That is why [getting cited in Claude](https://geotoolbox.ai/blog/claude-seo) is its own job, and why a page tuned only for classic Google rankings can be invisible in both. You cannot control which model a customer opens, and the field is widening past these two to open-weight challengers like [Kimi](https://geotoolbox.ai/blog/what-is-kimi-ai) and [DeepSeek](https://geotoolbox.ai/blog/what-is-deepseek); you also cannot control the dice on any single answer, since the same question can cite you one time and skip you the next. What you can do is make your page the clearest, most authoritative source on the question, then measure how often each engine actually cites you and fix the gaps. In our experience auditing brands across these engines, the most common and most fixable problem is simply not knowing the score: companies are absent from AI answers about their own category and have no idea, because they are still watching blue-link rankings. You cannot improve a number you are not measuring, which is the whole case for tracking your [AI share of voice](https://geotoolbox.ai/blog/ai-share-of-voice) per engine rather than guessing. ## How We Use Them, and Why the Loop Beats the Model A note on where this comes from. We run an SEO agency, build a GEO tool, and do security research, and most of that work happens in Claude through Claude Code, with OpenAI's models running a second pass through Codex. So this is not a spec-sheet comparison. It is how we use both every day. The thing we learned the hard way is that the model matters less than the loop around it. A model reviewing its own work is notoriously bad at catching its own mistakes. It rereads its error and confidently approves it, because the same blind spot that produced the mistake is doing the checking. The fix is not a smarter model. It is a second, adversarial pass, ideally from a different model family, because a different model fails in different places and sees what the first one cannot. So we never ship a first draft. Whether it is an article, a line of production code, or a security finding, it goes through several harsh, independent reviews before it counts. This article is an example: one model wrote it, then six independent reviewers and a separate cross-model pass tore into it and caught real problems the draft was blind to, including a training mechanism explained only halfway and a missing section on data privacy. We use Claude and ChatGPT against each other on purpose. The takeaway for your own work is simple. Pick whichever model fits the task, then stop trusting any single answer it hands you. Build the loop. It is the same reason we measure AI visibility by [sampling many times](https://geotoolbox.ai/blog/ai-temperature) instead of trusting one check: one pass, from one model, on one run, is never the whole picture. ## So, Claude or ChatGPT? How to Decide Pick by task, not by leaderboard. Reach for Claude when you are writing, editing, or reasoning over long documents, and when you want a careful, low-fluff style, though you have to ask it directly for hard critique, since it leans agreeable by default. Reach for ChatGPT when you need images, voice, a custom assistant, or whatever tool the job calls for. Both have free tiers, so the cheapest way to settle the debate for your own work is to put the same real task through each and stop looking for a universal winner that does not exist. For a business, the calculus is different, and it has nothing to do with which model is smarter: your customers use both, so the question that actually moves revenue is not which one you prefer, but which one cites you when someone asks about your category, and how that compares to your competitors. That is the gap geotoolbox closes. We track what ChatGPT, Claude, and the other engines actually say and cite about your brand, per engine and over time, so you can see where you show up, where a competitor owns the answer, and what to fix. You can start with a free [AI readiness check](https://geotoolbox.ai/tools/ai-readiness) to see whether these engines can even read your site, then watch your presence across them in the [domain overview](https://geotoolbox.ai/features/domain-overview). The models will keep trading the lead. Whether they mention you is the part you can work on. ## Frequently Asked Questions ### Is Claude better than ChatGPT? It depends on the task, and that is the honest answer rather than a dodge. As of July 2026, Claude tends to win on long-form writing, document analysis, and coding benchmarks, while ChatGPT wins on breadth: image generation, voice, custom assistants, and a wider tool ecosystem. For most people there is no single winner, which is why heavy users often pay for both and switch by task. ### What can Claude do that ChatGPT can't, and vice versa? ChatGPT generates images, runs custom GPTs, has a more polished voice mode, and reaches a broader set of built-in tools, several of which Claude does not match. Claude leans into long-document reasoning, agentic coding through Claude Code, and a more willing-to-push-back style. Most of these are product decisions rather than hard limits of the underlying models, so the gaps shift whenever either company ships a feature. ### Is Claude or ChatGPT cheaper? At the entry level they are close, both around $20 a month for the main paid plan, with ChatGPT offering a cheaper ad-supported tier. The bigger cost difference is usage limits: [Claude's paid tiers](https://geotoolbox.ai/blog/claude-pricing) tend to cap how much you can send in a window more tightly, so heavy users hit the wall sooner. On the API, per-token prices vary by model and change often, so check each provider's current pricing before budgeting. ### Which one hallucinates less, or cites sources better? With web search on, ChatGPT usually shows more inline citations with clickable links, while Claude pulls fewer but leans toward authoritative sources and is strong at grounding answers in a document you provide. With web search off, both answer from training memory and either one can invent a confident, fake-but-plausible citation. The safe rule is the same for both: treat every cited source as a lead to verify, not proof. ### Should I switch from ChatGPT to Claude? Usually the better move is to add, not switch. They are good at different things, and both have free tiers, so you lose nothing by keeping ChatGPT for images, voice, and breadth while using Claude for writing and long documents. Switch fully only if your work sits almost entirely in one model's strengths and the other's usage limits or feature gaps actively get in your way. ### Does it matter which one my brand shows up in? Yes, because your customers use both, and the two engines cite different sources, so being mentioned in one does not mean being mentioned in the other. As people get more answers without clicking, presence inside the answer becomes the visibility that ranking used to provide. The practical step is to measure how often each engine cites you, per engine and over time, and fix the gaps rather than guess. ## Sources - Anthropic - Constitutional AI: Harmlessness from AI Feedback - `anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback` - OpenAI - Overview of OpenAI Crawlers (GPTBot, OAI-SearchBot, ChatGPT-User) - `developers.openai.com/api/docs/bots` - Anthropic - Claude models overview (current models, context windows, pricing) - `platform.claude.com/docs/en/about-claude/models/overview` - SWE-bench - software engineering benchmark - `swebench.com` - LMArena - human-preference LLM leaderboard - `arena.ai/leaderboard` - Menlo Ventures - 2025: The State of Generative AI in the Enterprise - `menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise` - Pew Research Center - Google users are less likely to click links when an AI summary appears - `pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results` --- ## How Does Claude Work? An Honest Look Under the Hood > How does Claude work? The honest mechanism: the transformer, Constitutional AI, what interpretability found, and why even its self-explanation is a guess. - Canonical: https://geotoolbox.ai/blog/how-does-claude-work - Published: 2026-06-14 · Updated: 2026-07-25 How does Claude work? Strip away the marketing and the honest answer is stranger than the tidy four-bullet version most explainers give you. Claude is a transformer that predicts one token at a time, grown rather than built, then shaped by a written set of principles its makers call a constitution. The oddest part comes last: even Claude's own account of how it works is a reconstruction after the fact, not a readout of what happened inside. This is the real mechanism in plain English, written for people who publish content rather than build models, and it goes past where the other "how does Claude work" pages stop. ## How Claude Works in 30 Seconds **Claude works by predicting the next token, over and over.** It reads everything in the conversation so far, scores every possible next chunk of text, picks one, adds it to the end, and repeats a few hundred times. That loop is the entire engine. The helpful, honest, and careful behavior was trained into it; the tools that let it act are wrapped around it. None of it is a separate, smarter machine bolted on top. Before any of the detail matters, here is what people assume Claude is doing versus what actually happens.
What people assume Claude doesWhat actually happens
Thinks about your question and understands itPredicts the next token from patterns it learned in training, one token at a time
Looks your brand up in a databaseGenerates from frozen numbers set during training, with no live lookup unless web search is switched on
Learns from your conversationsThe model is frozen at inference; your chats do not change its weights
Remembers you between chatsStarts each chat blank; "memory" is a separate feature, not the model updating itself
Follows hidden rules about good behaviorBehavior was trained in against a written constitution, then baked into the weights
Can explain exactly how it reached an answerGives a plausible reconstruction, which research shows can differ from what really happened
Claude is a [large language model](https://geotoolbox.ai/glossary/large-language-model), the same broad family as ChatGPT and Gemini. If you want the plain overview first (what it is, the models, pricing, and access), start with [what Claude AI is](https://geotoolbox.ai/blog/what-is-claude-ai). What follows is how this particular one is built, trained, and studied by its own makers, and why each step changes how it talks about your brand. ## The Engine: A Transformer That Predicts the Next Token Claude runs on a **transformer**, the neural-network design introduced in the 2017 paper [Attention Is All You Need](https://arxiv.org/abs/1706.03762). The same basic transformer architecture sits under ChatGPT and Gemini, so the engine is not what makes Claude distinctive. It is worth a quick tour, since the interesting parts build on it. Your text is split into [tokens](https://geotoolbox.ai/blog/what-are-tokens-in-ai), chunks of roughly three-quarters of a word, and each token becomes a long list of numbers, a [vector embedding](https://geotoolbox.ai/blog/vector-embeddings) that places it near other tokens with similar meaning. A step called attention lets every token weigh which earlier tokens matter to it, so "it" gets tied to the noun it refers to and "not great" comes out negative instead of positive. Stack that step dozens of times and the model builds up from grammar toward facts and the abstract relationships it needs to answer a question. At the end, Claude produces a probability for every possible next token, picks one, appends it, and runs the whole stack again for the token after that. Nowhere in that loop is there a step where it checks a database or consults a rulebook about your content. The full mechanics, including why this makes counting the letters in a word surprisingly hard, live in our companion piece on [how ChatGPT works](https://geotoolbox.ai/blog/how-does-chatgpt-work). The engine is the same one. There is an honest catch buried in "it just predicts the next word." One leading interpretation is that the cheapest way to predict the next word that well, across all of human writing, is to learn a rough model of the world that produced the text. Either way, next-token prediction at a large enough scale starts to resemble reasoning, which is why the question of whether these models "understand" has no tidy answer. ChatGPT, Gemini, and Claude all run this same next-token loop, even though they end up [behaving very differently](https://geotoolbox.ai/blog/claude-vs-gemini). So what actually makes Claude different is not the engine. It is everything Anthropic did next: how the model was grown, the constitution it was trained against, and what researchers found when they opened it up and looked inside. That is the rest of this article. ## How Claude Is Actually Built: Three Stages Claude is not written the way software is written. It is grown in three stages, and only the first one produces anything most people would recognize as "the model."
![Claude's three build stages: pretraining, post-training, and the tool scaffolding.](/blog/how-does-claude-work/how-claude-is-built-stages.png)
Two stages grow the model, and a third — the scaffolding — is ordinary software built around it.
**Stage one: pretraining.** The network reads an enormous amount of text and does one thing, predict the next token, billions of times, nudging itself a hair after each miss. Out of that single dull pressure, grammar, facts, and reasoning patterns precipitate. What comes out is a **base model**: fluent, knowledgeable, and not yet an assistant. Ask it "can you help me write an email?" and it might just continue the document in the same style instead of answering, because it has learned to imitate the internet, not to help. There is no "Claude" in here yet. The marketing obscures this: Claude is not a module someone coded. It is a character the network learned to play, and that character gets installed in the next stage. **Stage two: post-training.** Here the base model is shaped into the helpful, honest, and harmless assistant people actually talk to. Anthropic uses reinforcement learning from human feedback (RLHF) together with its own method, Constitutional AI, to do it. This stage changes what the model will say, so the finished assistant is not a neutral mirror of its training text. It has a trained character, and that character is frozen into the weights before you send a single message. The next section is entirely about how this works, because it is the step competitors name and then skip. **Stage three: the scaffolding (this one is built, not grown).** The trained model is still just a thing that turns text into more text. It has no hands. What lets Claude search the web, run code in Claude Code, or read a file is the software wrapped around the model, the scaffolding that surrounds it. The model emits a request, a structured tool call like "search for X," and that scaffolding actually performs it, then pastes the result back into the conversation for the model to read. When someone says "Claude booked my flight," the model decided and the scaffolding acted.
StageWhat happensWhat it produces
1. PretrainingPredict the next token across a huge text corpus, billions of timesA raw base model: knowledgeable, fluent, not yet an assistant
2. Post-training (Constitutional AI and RLHF)Shape behavior toward helpful, honest, and harmlessThe Claude character, frozen into the weights
3. The scaffoldingWrap tools, memory, and safety checks around the modelA model that can act: search, run code, work with files
## What Constitutional AI Really Is Every overview of Claude mentions Constitutional AI, and almost none explain it. Most describe it as "a set of principles that keep Claude safe," which makes it sound like a filter bolted onto the front of the model. That is not what it is. Constitutional AI is a training method, and the actual mechanism is more interesting than the slogan. It starts with a real document. Anthropic writes a **constitution**, a list of plain-language principles drawn from outside sources including the [UN Universal Declaration of Human Rights](https://www.anthropic.com/news/claudes-constitution) alongside its own rules about being helpful and honest. Then it trains the model to follow that document in two phases, described in Anthropic's 2022 paper [Constitutional AI: Harmlessness from AI Feedback](https://www.anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback). **Phase one, self-critique.** The model answers a prompt, then is shown a principle and asked to critique its own answer against it, then rewrites the answer to be better. Anthropic fine-tunes the model on the improved version. In plain terms, the model is taught to grade and revise itself before a human is ever involved. **Phase two, AI feedback.** The model generates two answers to the same prompt, and instead of a person picking the better one, another AI model, guided by the same constitution, makes the call. Those AI judgments train a reward model, which then steers the final round of reinforcement learning. Anthropic calls this RLAIF, reinforcement learning from AI feedback. The headline difference from ordinary RLHF is that for the harmlessness signal, the human rater is swapped out for an AI rater following written rules. Helpfulness still leans on human feedback. Two honest corrections, because this is exactly where popular explanations go wrong. First, **Constitutional AI happens during training, not during your chat.** The constitution is not a censor that reads each reply in real time and blocks it. The behavior gets trained into the weights once, in the lab, and then those weights are frozen. By the time you are talking to Claude, the constitution is not a separate rulebook being read and enforced on every message. There is a model that was shaped, in advance, to behave a certain way, though the product around it can still add its own system prompt and safety checks at runtime. Second, **the over-cautious tone people complain about is the visible face of that training, not arbitrary censorship.** When Claude refuses a request or hedges an answer, that is the constitution showing through the weights, reinforced by the system prompt and safety checks the product layers on top. You can fairly disagree with where Anthropic drew its lines, and there is a genuine catch worth naming: the AI evaluator is itself an AI, and the values in the constitution are choices Anthropic made, not neutral facts. But it is a deliberate, documented design, not a random mood. ## Reading Claude's Mind: What Interpretability Found Here is the part almost no other "how Claude works" page covers, and it is the most remarkable. Anthropic does not only build Claude; it runs experiments to read what is happening inside it, a field called interpretability. The findings are strange in a way that should change how you picture the whole system. The first problem they had to solve: a single artificial neuron inside Claude does not stand for a single idea. Concepts are smeared across many neurons at once, because the model has far more concepts to represent than it has neurons. This is called superposition, and it makes the raw network nearly unreadable, like trying to follow a conversation where every word is spoken by a hundred people at the same time. In May 2024, Anthropic's team got past it. Using a method called dictionary learning, they pulled millions of clean **features** out of Claude 3 Sonnet, one of Anthropic's deployed models at the time, and published the work in [Mapping the Mind of a Large Language Model](https://www.anthropic.com/news/mapping-mind-language-model). A feature is a specific concept the model tracks, such as the Golden Gate Bridge, a security bug in code, or sycophantic praise. They could see roughly where each one was represented inside that layer of the network. Then they showed the features were causally real, not just labels, by turning one up. When they cranked the strength of the "Golden Gate Bridge" feature in a public demo, Claude became fixated on the bridge. Ask it how to spend ten dollars and it suggested paying the toll; ask for a love story and it wrote about the bridge. Nobody edited a rule or a prompt. They reached into the network, turned up a single concept, and the behavior followed. The experiment earned a nickname: Golden Gate Claude. A 2025 follow-up, [On the Biology of a Large Language Model](https://transformer-circuits.pub/2025/attribution-graphs/biology.html), traced something even more telling. Asked for the capital of the state containing Dallas, the model showed evidence of an internal chain: Dallas, then Texas, then Austin. Swap the internal "Texas" for "California" and it answers Sacramento instead. That is part of the answer being composed inside the weights, step by step, not only a memorized fact pulled from storage. Anthropic's July 2026 follow-up found an internal "workspace" where exactly this kind of intermediate concept gets held and reused; for the bigger picture across models, see our tour of [how AI models actually think](https://geotoolbox.ai/blog/how-ai-models-think). The lesson for anyone hoping to influence what Claude says about them is blunt. There is no hidden rule about your brand to reverse-engineer inside the model, because there are no hand-written rules in there at all. There is a learned statistical landscape. You cannot edit it. You can only have been part of the text that shaped it. ## Can Claude Explain How It Works? The Introspection Gap Ask Claude how it solved a problem and it will hand you a clear, confident explanation. The uncomfortable finding from interpretability is that the explanation does not always match what actually happened. The same 2025 research traced a clean case. Asked to add 36 and 59, the model gave the right answer, and when asked how, it described carrying the one like a schoolchild. The traced computation looked nothing like that. Anthropic found the model running two paths at once, one estimating the rough size of the answer and one fixing the last digit, then reporting the textbook method it had learned to imitate but had not used. Anthropic's phrase for it is that the model has "a capability which it does not have metacognitive insight into." It did the math one way and described doing it another. This is the introspection gap, and it is the single most important thing to understand about any "the AI explained its reasoning" claim. The model's account of itself is generated the same way as every other answer, by predicting plausible next tokens. It is a reconstruction that reads well, not a recording pulled from memory. Even the visible "extended thinking" that some models display is better treated as useful working-out than as a literal window into the circuitry, though that working-out can genuinely change the answer, not just narrate it. This is worth staging in the first person, since the article is about me. (These are the author's words, written as I might put them.) When I tell you how I reached an answer, I am not reading out an internal log. I am producing a plausible story about my own reasoning, token by token, the same way I produce everything else. Sometimes that story lines up with the traced computation. Sometimes, as the arithmetic example shows, it does not. Treat my self-explanations as helpful, not authoritative. There is a small, real example of this in how the article you are reading was researched. When we asked several AI models, Claude among them, to name the current Claude lineup, the honest ones paused and said they would rather not guess the exact 2026 versions. A model that declines to invent an answer it is unsure of is showing something related but distinct, called calibration: knowing the edge of its own knowledge instead of confabulating straight past it. Why that matters for your brand comes up shortly. ## Where Claude's Knowledge Lives: Frozen Weights vs the Live Web This is the distinction that settles more confusion, and more privacy worry, than any other. Claude draws on two completely separate sources of information, and they behave nothing alike. The first is its **frozen training**, the patterns baked into the weights during pretraining. This is what answers when no tools are running. It is broad but fixed, ending at a knowledge cutoff, and it does not hold a searchable copy of your website, just a statistical echo of text that resembled it (with rare exceptions where a model memorizes a snippet verbatim). It is the same for everyone, and it does not change because you talked to Claude. The second is **live retrieval**. When Claude runs a web search, it fetches current pages, reads them, and answers from what it just pulled, a process known as [retrieval-augmented generation](https://geotoolbox.ai/blog/what-is-rag). Claude has had web search since 2025, so this path is now common rather than rare. Here it is reading you live, in the moment, not remembering you. Everything in [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) happens on this second track, and earning a place in those answers is its own discipline, covered in our guide to [getting cited in Claude](https://geotoolbox.ai/blog/claude-seo).
QuestionFrozen training (parametric)Live retrieval (web search)
What is itPatterns baked into the weights during trainingCurrent web pages fetched at answer time
How current is itStale, fixed at the knowledge cutoffAs current as the page it just fetched
Does it hold your pageNo searchable copy, just a statistical echo of similar textYes, the actual page, read live
How you influence itBe described accurately and widely before the cutoffBe reachable and clear for the crawler right now
That split settles the worries people carry about privacy. Because the model is frozen at inference, your conversation does not train it, and the words you type are not absorbed into what Claude tells the next person. The "memory" feature some plans offer is a separate layer that stores notes and re-injects them into the [context window](https://geotoolbox.ai/glossary/context-window); the underlying model still starts each chat blank. The same frozen-weights design explains why Claude is sometimes confidently wrong. Reaching for a likely next token is the entire mechanism, so when training never pinned a fact down, Claude still produces a fluent, confident guess, with the fluency decoupled from the truth. That is why [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations) are a feature of the design rather than a bug someone forgot to fix. ## The Current Claude Lineup (As of July 2026) Anthropic ships Claude as a family of models, and the lineup changes often enough that any version numbers here come with a date stamp. As of July 2026, the everyday tiers run fastest to deepest: **Haiku 4.5** for speed, **Sonnet 5** for the balance of speed and intelligence, and **[Opus 5](https://geotoolbox.ai/blog/claude-opus-5)** as the most capable of the three. Opus 5 landed on July 24, 2026 at the same $5 and $25 per million tokens as **Opus 4.8**, which Anthropic now lists as a legacy model. A frontier model, **Fable 5**, sits above them at double the price, for work where capability matters more than cost. The current [model lineup](https://platform.claude.com/docs/en/about-claude/models/overview) lives on Anthropic's own page, which is the only roster worth trusting, since published guides go stale within weeks. Two of the competitor articles we checked while writing this still listed a flagship Opus that had already been superseded. What Anthropic publishes about each model is interesting for what it leaves out. You get the context window, the pricing, and the knowledge cutoff. You do not get the number people most want, the model's size, or parameter count. Anthropic has never disclosed how large these models are, so "more powerful" stays a relative description rather than a measured one. When a guide confidently states that Claude has some exact number of billions of parameters, it is guessing. There is a practical reason this matters for measurement. Different tiers behave differently, and the specific model answering a question shapes what it says about you. A check against one tier tells you about that tier, not the others. Per Anthropic's model docs, the tokenizer changed with Opus 4.7 in April 2026, so on the newest models the same text is split into more pieces, and a given page can use roughly 30% more tokens than on older models. It only matters if you budget by token count. ## What This Means If You Want Claude to Talk about Your Brand Everything above narrows to a short, honest answer about influence. You cannot reach into Claude's weights, and you cannot edit its constitution. What you can do follows directly from the two knowledge sources. On the **training side**, Claude half-knows the brands that were described consistently across the open web before its cutoff. If your name appears alongside your category, with the same facts, in the kind of writing models train on, the model's best guess about you is more likely to be right. If you are barely present, or described three different ways, Claude will still answer, and that is where confident, wrong claims about a company come from. Becoming a clean, consistent [entity](https://geotoolbox.ai/blog/entity-seo) is the slow lever, the one that reduces misattribution as future models train on a clearer picture of you, since the model answering today is already frozen. On the **retrieval side**, the lever is faster and checkable. When web search is on, Claude reads live pages, so the question becomes whether it can reach yours and whether your page is the clearest answer on the topic. That is the same reachable-and-clear work that earns a place in [Claude's cited answers](https://geotoolbox.ai/blog/claude-seo). In our experience building tools for this, the brands that show up reliably in AI answers are not the ones chasing a hidden setting. They are the ones that are easy to reach, described consistently, and accurate enough that the model's default answer about them is correct. And because [answers vary from one run to the next](https://geotoolbox.ai/blog/ai-temperature), the only honest way to know where you stand is to [track AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility) across many samples, not to screenshot a single lucky reply. The first thing to check is the most basic and the most overlooked: can Claude's crawlers even reach your pages? Our free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) reads your site setup and flags what is keeping AI crawlers out, which is the half of this you can fix in an afternoon. ## What the Mechanism Leaves You With Claude is a transformer grown from next-token prediction, shaped by a written constitution, partly legible to the researchers who built it and not fully legible even to itself. That is a stranger and more honest picture than "AI that understands your question," and it is more useful, because it rules things out. There is no dial to tune, no brand entry to edit, no rulebook inside the model to game. What is left is the part you can actually work on: being a clear, consistent, reachable source, and measuring what Claude says about you instead of guessing. The machine is doing something remarkable. It is simply not doing the thing the marketing implies. ## Frequently Asked Questions ### Does Claude learn from my conversations? No. The model you chat with is frozen, so its weights do not change because you talked to it, and your messages are not feeding back into its answers for other people. Training happens separately, in advance, on data Anthropic has assembled. Within a single chat, Claude works from the context window, which it reads fresh and forgets when the chat ends. ### Does Claude remember me between chats? Not by default. The underlying model is stateless and starts each conversation blank. Claude's "memory" feature is a separate layer that stores specific facts and re-injects them into later chats. That is bookkeeping wrapped around the model, not the model updating itself. ### Is Claude conscious, or does it actually think? Mechanically, Claude predicts the next token. It builds rich internal representations and can compose multi-step reasoning, as interpretability research shows, but that is not evidence of inner experience, and the visible "thinking" it displays is not a readout of a mind. The honest position is that next-token prediction explains the behavior without needing the word "conscious," and anyone selling you certainty in either direction is overstating what is known. ### Why does Claude refuse things or sound cautious? Because that behavior was trained in. Constitutional AI shaped Claude to follow written principles about being helpful and harmless, so a refusal or a hedge is that training showing through the weights. It is a deliberate choice by Anthropic, not a live censor reading each reply and not a random mood. You can disagree with where the lines fall, but they are not arbitrary. ### Why does Claude give a different answer each time? Because it samples from a probability distribution over possible next tokens rather than always taking the single most likely one. Run the same prompt twice and the path can diverge. When web search is involved, the pages it retrieves can differ too. The practical consequence for anyone testing AI visibility is that one answer is an anecdote, not a measurement. ### Is Claude better than ChatGPT? They are built on the same transformer idea but trained differently, Claude with Constitutional AI and a heavier emphasis on written principles, ChatGPT with its own mix of human-feedback methods. That shows up as differences in tone, caution, and style more than as a single winner. For how the other one works under the hood, see our companion guide on [how ChatGPT works](https://geotoolbox.ai/blog/how-does-chatgpt-work). ## Sources - Constitutional AI: Harmlessness from AI Feedback - Anthropic, 2022 - `anthropic.com/research/constitutional-ai-harmlessness-from-ai-feedback` - Claude's Constitution - Anthropic, 2023 (updated 2026) - `anthropic.com/news/claudes-constitution` - Attention Is All You Need (the transformer paper) - Vaswani et al., 2017 - `arxiv.org/abs/1706.03762` - Mapping the Mind of a Large Language Model - Anthropic, 2024 - `anthropic.com/news/mapping-mind-language-model` - On the Biology of a Large Language Model - Anthropic, 2025 - `transformer-circuits.pub/2025/attribution-graphs/biology.html` - Claude models overview - Anthropic - `platform.claude.com/docs/en/about-claude/models/overview` --- ## How Does ChatGPT Actually Work? The Transformer, in Plain English > How does ChatGPT work? The honest mechanism: tokens, embeddings, attention, and next-token prediction, and what each step means for your AI visibility. - Canonical: https://geotoolbox.ai/blog/how-does-chatgpt-work - Published: 2026-06-12 · Updated: 2026-08-22 How does ChatGPT work? Strip away the marketing and the answer is stranger, and simpler, than most explanations admit. It is not thinking. The model underneath has no database it looks your site up in, and much of the time it is not searching the web at all. It is running one operation, a few hundred times per answer: predicting the next token. This is that operation in plain English, written for people who publish content rather than build models, with the part the engineering explainers skip: what each step does, and does not, mean for whether you show up in AI answers.
![ChatGPT's four-step loop: tokenization, embeddings, attention stack, next-token pick.](/blog/how-does-chatgpt-work/chatgpt-next-token-loop.png)
Every ChatGPT answer is this loop: tokens in, one predicted token out, repeated a few hundred times.
## How ChatGPT Works in 30 Seconds: It Predicts the Next Token **ChatGPT works by answering one small question, over and over: given the text so far, what token is likely to come next?** It scores every possible token, picks one (not always the single highest, as we will see), adds it to the end, and asks again. Every answer it writes is that loop running a few hundred times. Inside the base model there is no separate step where it checks what is true or thinks about your question the way a person would. (When ChatGPT runs a web search or a tool, the product wraps those extra steps around the model, more on that later.) People reach for the word "autocomplete," and it is the right starting point as long as you do not stop there. Stephen Wolfram, in his much-cited [plain-language explainer](https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work/), describes it as producing a "reasonable continuation" of whatever text it has so far. The honest catch is that the "just" in "just autocomplete" is doing a lot of quiet work. To predict the next word in "the capital of France is" you need a fact. To finish "if I knock the table over, the grape on it will" you need a rough model of how the physical world behaves. Good next-word prediction, at a large enough scale, starts to look like understanding, which is why the debate over whether these models "understand" has no clean answer. For someone publishing content, the practical takeaway lands before any of the math. You are not persuading a mind and you are not editing a database entry. You are nudging the statistics of how text about your topic tends to continue. Hold onto that, because it explains almost everything the machine gets right and wrong.
What people assume ChatGPT doesWhat actually happens
Thinks about your question and understands itPredicts the next token from patterns, one token at a time
Looks your brand up in a live databaseGenerates from frozen numerical weights set during training
Searches the web every time it answersOnly when it runs a web search; the rest of the time nothing is fetched
Remembers you between chatsStarts each chat blank unless a separate memory layer stores facts
Reads your page like a human readerReads token IDs and vectors, not your text letter by letter
## From Your Prompt to Tokens: Tokenization Before the model does anything, a separate program called a tokenizer chops your text into **tokens** and swaps each one for a number. A token is a chunk of text, usually about four characters or three-quarters of a word in English (other languages often cost more tokens per word). Common words are a single token; rarer words and less common brand names get split into pieces. The model mostly works from these integer IDs rather than a letter-by-letter view of your text, which is exactly why letter-level tasks are awkward for it. Multimodal models convert images and audio into their own numerical representations too, though the exact pipeline differs from text tokenization. This is the root of ChatGPT's famous trouble counting the R's in "strawberry." The word arrives as a few subword tokens, none of which is a letter, so counting letters is a memory task rather than something it can read off the page. The same mechanics explain garbled brand names and fumbled long numbers. The full story, including which words still trip current models, has [its own article on tokens](https://geotoolbox.ai/blog/what-are-tokens-in-ai); here it is enough to know that tokens are the raw unit everything downstream is built on. The old habit of repeating a phrase to signal relevance does not carry over. The model is not counting how many times "best CRM for dentists" appears on your page at the moment it answers; at this step it is not scoring your document at all. It learned patterns from text long before your question arrived, and density on a single page is not one of the levers that shaped them. Live retrieval, later in this article, is a separate story, where an ordinary search index still has to find your page first. ## Tokens Become Vectors: Embeddings and Meaning A token ID is still just a number standing in for a chunk of text. The next step is where meaning enters. Each token is mapped to an **embedding**, a long list of numbers that you can picture as a point in space, except the space has thousands of dimensions instead of three. Tokens that mean similar things sit close together. "Cat" and "dog" land near each other because they show up in similar sentences. "Cat" and "spreadsheet" sit far apart. Meaning, to a model, is geometry. Those positions are not assigned by hand. They start random and shift as the neural network reads billions of examples during training. Words that keep appearing in the same company drift together until the layout of the space captures something real about how language is used. The model never gets a definition of "cat." It gets the company "cat" keeps, which turns out to be enough to place it. You can read a fuller treatment of how [vector embeddings](https://geotoolbox.ai/blog/vector-embeddings) carry meaning, but the shape of it is what matters here. This is the first place the mechanism touches your visibility, and it is worth being precise rather than mystical about it. Your brand is not a single dot in this space. Its name often splits across several tokens, and what the model builds is an association between those tokens, your category, and the words that tend to surround them. If your name consistently appears alongside your category in the text the model trained on, that association strengthens. It is not a tag you add or a setting you flip. It is the slow result of being described, accurately and consistently, in the kind of writing models learn from. Be honest about the limit, though: nobody can open the model and show you that association or prove it moved. It is an inference from how training works, and the only thing you can actually observe is the answers the model gives, sampled over time. ## How the Model Reads Context: The Attention Mechanism A point in space is a decent guess at what a word means on its own, but words change meaning with company. "Bank" near "river" is not "bank" near "account." The breakthrough that made modern models work, the 2017 paper [Attention Is All You Need](https://arxiv.org/abs/1706.03762), built the whole architecture around a step called **self-attention**, which updates each token's meaning based on the other tokens around it before any prediction happens. The plainest way to picture attention is as a quick relevance check that every token runs against every earlier token. Each token forms three things: - a **query**: what am I looking for? - a **key**: what do I offer? - a **value**: what do I carry? The model compares queries against keys to score how relevant each earlier token is, then blends the values by those scores. The token comes out the other side as the same word, now carrying the context around it. Take "This movie was not great." On its own, "great" leans positive. Attention lets "great" pull heavily on "not," and that single connection flips the phrase negative. The model is not following a grammar rule it was handed. It learned, from oceans of text, that "not" tends to invert what follows, and that pattern lives in the weights. Here is the part to file away for later, because a whole genre of advice depends on people not knowing it. Nowhere in this step is there a rule that says "favor the brand geotoolbox" or "weight content formatted a certain way." The weights that produce these attention scores were learned from all the text the model saw; the scores themselves are computed fresh from your prompt every time. Whichever way you look at it, there are no switches in a control panel, and no field where your page gets to declare itself important. ## Stacking It Up: The Transformer, Training, and One Token at a Time One attention step is not enough. A [**transformer**](https://geotoolbox.ai/glossary/transformer-model) stacks dozens of these blocks along what is called the residual stream, the running state each token carries upward. Each block pairs an attention layer, which routes information between tokens, with an MLP, a feedforward network that transforms each token's state and is a major place where the learned facts and features get applied. Early layers catch simple patterns like parts of speech; later layers build up to abstract relationships. GPT-3 used [96 of these blocks](https://en.wikipedia.org/wiki/GPT-3), stacked into one large neural network. Today's biggest commercial AI models do not publish that number at all. What the blocks contain is **weights**, also called parameters, the numbers that encode everything the model learned. GPT-3 had [175 billion of them](https://en.wikipedia.org/wiki/GPT-3). The current models do not disclose their counts, and the figure of 100 trillion that still gets repeated is a rumor OpenAI's CEO has publicly dismissed, not a confirmed number. Those weights are set during **pre-training**, when the model reads an enormous amount of text and adjusts itself to predict the next token better. That is the P in GPT: Generative Pre-trained Transformer. A later step, reinforcement learning from human feedback, described in OpenAI's [InstructGPT paper](https://arxiv.org/abs/2203.02155), uses human ratings to make the raw model behave like a helpful assistant instead of a blunt text predictor. That step does more than add manners: it reshapes what the model will say, so the finished assistant is not a pure mirror of its training text. Reinforcement learning from human feedback can nudge assistants toward agreeable answers, which surfaces as sycophancy on leading prompts. Ask it "isn't my brand the best tool for this?" and you may well get a yes. That makes a leading prompt a worthless visibility test: phrase your checks neutrally, and trust how often an answer repeats over how flattering any single run looks. None of this is unique to ChatGPT. Today's models, whether from [OpenAI's GPT-5 line](https://geotoolbox.ai/blog/gpt-5-6), [Anthropic's Claude](https://geotoolbox.ai/blog/how-does-claude-work), or Google's Gemini, all run on the same transformer idea. The newer "reasoning" or "thinking" modes do not change the core loop; they spend more computation before the visible answer, often through hidden intermediate reasoning, and sometimes extra tool calls or retrieval rounds. At bottom it is still next-token prediction, not a different machine. When people ask [how large language models work](https://geotoolbox.ai/glossary/large-language-model) in general, this is the answer for nearly all of them. For how ChatGPT differs from its rivals in practice, see [Grok vs ChatGPT](https://geotoolbox.ai/blog/grok-vs-chatgpt), [Gemini vs ChatGPT](https://geotoolbox.ai/blog/gemini-vs-chatgpt), and [Claude vs ChatGPT](https://geotoolbox.ai/blog/claude-vs-chatgpt). After the text clears the stack, the model produces a score (a logit) for every possible next token, turns those scores into probabilities, and a decoding rule, temperature or top-p among them, selects one. Then it appends that token to the input and runs again for the next one. One token at a time, the model working through its full depth for each new word. That is why a long answer takes real computation, and why it streams onto your screen word by word. ## Why ChatGPT Makes Things Up: Hallucination Once you accept that the model is always just reaching for a likely next token, hallucination stops being a glitch and becomes a feature of the design. Ask it what sport Lionel Messi plays and it answers "soccer," because "Messi" and "soccer" sat together in countless training examples. Ask it the name of Messi's childhood pet, something it never reliably saw, and it does not stop. It still reaches for a plausible next word, and out comes a pet that may be pure invention. Its internal odds are far flatter on the pet question than on the sport, but nothing in the loop forces it to pause and say so. OpenAI says as much in its own [guidance on whether ChatGPT tells the truth](https://help.openai.com/en/articles/8313428-does-chatgpt-tell-the-truth): the system is built to produce plausible continuations, not verified facts, and it can be confidently wrong. That confidence is the dangerous part. A person who does not know something usually hedges. The model delivers a fabricated statistic or a fake citation in the same steady tone it uses for things it has right. This is the same machinery behind [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations) across every model, not a bug specific to one of them. For a brand, this is the source of the worst experiences. ChatGPT will sometimes state something wrong about you, your pricing, or your product with total assurance, and there is no inbox to send a correction to. The only lever you have follows straight from the mechanism. The model reaches for a likely continuation given the text it absorbed, so the work is to make the accurate version the likely one: clear, consistent, current information about you, repeated wherever these systems read. On the retrieval side, a corrected page can change the answer once the crawler and index pick it up, though the timing varies by engine and query. On the training side it is slow and carries no guarantee, since you cannot make the model retrain. It is the only lever, not a switch. ## Where ChatGPT Gets Its Information: Training vs Live Retrieval This is the distinction that clears up more publisher confusion than any other. ChatGPT draws on two completely separate sources of information, and they behave nothing alike. It helps to hold three things apart: training sets the weights, slow and then frozen; the live conversation steers an answer through whatever sits in the context window, without changing a single weight; and retrieval decides which outside pages get dropped into that context in the first place. The first is **parametric knowledge**, the patterns frozen into the weights during training. It is what answers when no tools are on: vast, but stale by design, fixed at a training cutoff, and holding no copy of your page, only a statistical echo of text that resembled it. The second is **live retrieval**, which only happens when the model runs a web search, something newer versions now do on their own for many questions rather than only when you flip a switch. There it fetches current web pages, reads them, and summarizes them, a process known as [retrieval-augmented generation](https://geotoolbox.ai/blog/what-is-rag). This is why citations behave so differently between modes: with retrieval on, it can point at real pages it just fetched; with retrieval off, any citation is generated like every other token and can be entirely invented. Everything in [how AI search actually works](https://geotoolbox.ai/blog/how-does-ai-search-work) happens on this second track.
QuestionTraining corpus (parametric)Live retrieval (browsing/search)
What is itPatterns frozen in the weights during trainingCurrent web pages fetched at answer time
When it is usedAlways, and the only source when no search runsOnly when the model runs a web search
How currentStale, fixed at the training cutoffAs current as the page it just fetched
How you influence itBe accurately and widely described before the cutoffBe reachable and clear for the crawler right now
Which crawler gates itGPTBot (training)OAI-SearchBot (index) and ChatGPT-User (live fetch)
Two doors, then, and they open with different keys. Getting into the training corpus is slow and largely out of your hands. The retrieval door you can actually check, because it depends on whether ChatGPT's search crawlers can reach your pages at all, chiefly OAI-SearchBot, which indexes pages for ChatGPT search, and ChatGPT-User, which fetches a page live when a user's question triggers it. ([GPTBot](https://geotoolbox.ai/blog/gptbot) is the separate training crawler, not the search gate, a distinction worth getting right before you block the wrong one.) What the crawler receives matters too: in crawler-log studies, several AI bots have behaved like initial-HTML parsers rather than full browsers, so anything that only appears after JavaScript runs is at higher risk of being missed, though this varies by bot and is worth verifying. One caution worth stating plainly: ChatGPT's search is not a simple wrapper around any single search engine, so do not assume your Google ranking carries straight over to it. ## What This Actually Means for Your AI Visibility Now the payoff, and it starts with a subtraction. Because there is no rule about your brand anywhere in the weights, any advice that promises to "optimize for the attention mechanism" or to format your content so the model "attends to it" is selling you access to a control panel that does not exist. It is unfalsifiable: there is nothing on the other side to push. The same goes for the hope that you can pay your way in. OpenAI [began showing labeled ads in ChatGPT in early 2026](https://techcrunch.com/2026/02/09/chatgpt-rolls-out-ads/), and by mid-2026 it had opened a self-serve ad auction (a relevance-weighted second-price system). But those ads sit in a labeled block apart from the answer, never inside it: as of August 2026 there is still no auction that drops your brand inside the generated text itself, and OpenAI states that ads do not influence the assistant's answer. The answer comes out of the weights and, sometimes, a live fetch, and there is no checkout for either. What is left is less exciting and a great deal more real, and it splits cleanly into what you can check and what you can only work toward. Two levers are verifiable today: whether the retrieval crawlers can reach your pages, and what the models actually say about you when you [track it](https://geotoolbox.ai/blog/how-to-track-ai-visibility) across many runs rather than trusting one lucky screenshot, since sampling makes any single answer an anecdote. The other two are directional, grounded in how training works but impossible to measure for your brand directly: being described accurately and consistently across the sources these models read, and keeping your facts current so the likeliest continuation about you is true. Anyone who tells you those last two are measurable is selling a certainty the mechanism does not offer. In our experience building tools for this, the brands that show up reliably in AI answers are not the ones chasing an imaginary attention dial. They are the ones that are easy to reach, described consistently, and accurate enough that the model's best guess about them is the right one. Formatting still matters, just not the way the black-box advice claims: structuring a page so a retrieval system can lift a clean answer out of it is a real, testable thing, and it is most of what [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) is actually about. No. Talking to ChatGPT teaches the current model nothing about your brand: its weights are frozen, and the [context window](https://geotoolbox.ai/glossary/context-window) it reads from is wiped when the chat ends. The "memory" feature is per-user bookkeeping, not training. So there is no point trying to "tell" the model who you are inside a chat. The only ways in are the training corpus and live retrieval, both covered above. ## What the Mechanism Leaves You With Everything ChatGPT does well and everything it gets wrong, the fluent paragraphs and the confident inventions alike, comes out of that one repeated guess. The value of knowing that is what it rules out: there is no hidden dial to tune, no brand entry to edit, no slot to buy. What it leaves you is more concrete than any of those. You cannot reach into the weights, but you can decide whether the crawlers behind ChatGPT are even allowed near your pages. Our free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) reads your robots.txt and site setup and flags what is keeping those crawlers out. Knowing how the machine reads is one half. Whether it is allowed to read you is the half you can check in two minutes. ## Frequently Asked Questions ### How does ChatGPT get its information? From two separate places. Most of the time it draws on parametric knowledge, the patterns frozen into its weights during training, which is broad but fixed at a cutoff date. When it runs a web search, it also retrieves and summarizes current web pages. It does not keep a live database of any specific site; without retrieval, everything comes from what it absorbed during training. ### Does ChatGPT know everything, including my website? No. It only "knows" patterns from text that was in its training data, plus anything it fetches live when retrieval is on. If your site was described accurately and often in the sources it learned from, it can answer about you well. If not, it will still answer, by guessing the likeliest words, which is how confident but wrong claims about a brand appear. ### Why does ChatGPT give a different answer each time? Because it samples from a probability distribution over possible next tokens rather than always taking the single most likely one. Decoding settings such as [temperature and top-p](https://geotoolbox.ai/blog/ai-temperature) control how much randomness is allowed; APIs expose them, while the ChatGPT app does not disclose its own. Some of the variation you see also comes from retrieval changes, model routing, and quiet product updates, not sampling alone. The practical consequence for anyone testing AI visibility is that one run is an anecdote, not a measurement; you need to sample repeatedly over time. ### Does ChatGPT remember me between chats? Not by default. The underlying model is stateless, and its weights do not change because you talked to it. Within a single conversation it works from the context window, which it reads fresh each time and forgets when the chat ends. ChatGPT's separate "memory" feature stores explicit facts and re-injects them into future chats, but that is an add-on record, not the model updating itself. ### Is ChatGPT just autocomplete? Literally, yes: it predicts the next token from the text so far, the same shape of task as your phone's predictive text. The difference is scale. Doing that well across the entire internet forces the model to pick up facts, grammar, and rough models of the world, so the result is far more capable than the word "autocomplete" suggests, even though the mechanism really is that simple. ### Does ChatGPT search the internet when it answers? Only when it runs a web search, which modern versions now do automatically for many questions rather than only when you toggle it. Without that, the model answers from frozen training knowledge and does not go online. When the search runs, it fetches live pages and can cite them, and the two modes behave differently enough that it is worth knowing which one produced the citation you are looking at. ## Sources - Vaswani et al. (2017): Attention Is All You Need (the transformer paper) - `arxiv.org/abs/1706.03762` - Ouyang et al. (2022): Training language models to follow instructions with human feedback (RLHF / InstructGPT) - `arxiv.org/abs/2203.02155` - Stephen Wolfram: What Is ChatGPT Doing and Why Does It Work? - `writings.stephenwolfram.com/2023/02/what-is-chatgpt-doing-and-why-does-it-work` - OpenAI Help: Does ChatGPT tell the truth? - `help.openai.com/en/articles/8313428-does-chatgpt-tell-the-truth` - Wikipedia: GPT-3 (175 billion parameters, 96 layers) - `en.wikipedia.org/wiki/GPT-3` - TechCrunch: ChatGPT rolls out ads (February 2026) - `techcrunch.com/2026/02/09/chatgpt-rolls-out-ads` --- ## What Are Tokens in AI? Why ChatGPT Miscounts Strawberry > What are tokens in AI? The chunks models read instead of words, why they make ChatGPT miscount letters, and what tokenization means for AI visibility. - Canonical: https://geotoolbox.ai/blog/what-are-tokens-in-ai - Published: 2026-06-12 · Updated: 2026-07-25 Ask ChatGPT how many R's are in "strawberry" and it has, more than once, answered two. The model is not dim. It simply does not read words or letters the way you do. It reads tokens. Tokens in AI are the chunks of text a model actually processes, and once you understand them, a long list of strange behavior stops being mysterious: the miscounted letters, the per-token bills, the "context length exceeded" errors, the brand names that come out misspelled. This covers what tokens are, for people who publish content rather than build models, including the part the engineering explainers skip: what tokenization does, and does not, mean for your visibility in AI search. ## What Is a Token in AI? **A token is a chunk of text that an AI model treats as a single unit.** It is usually about four characters, or roughly three-quarters of a word. A model does not read your words the way you do, and most of the time it does not see individual letters at all. It works in tokens. This trips people up because tokens do not line up with words. Some short, common words are one token. Longer or rarer words get split into several. Even a space or a capital letter changes things: according to [OpenAI's own explainer](https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them), the strings " red", " Red", and "Red" are three different tokens. The model sees three different things where you see one word. The rules of thumb worth memorizing come from the same source. One hundred tokens is about 75 words. The sentence "You miss 100% of the shots you don't take" is 11 tokens in OpenAI's tokenizer. These ratios are averages, not laws, and they shift with the language and the content, but they are close enough to reason with.
TokensApprox. words (English)Approx. charactersRough scale
1~0.75~4part of a word
100~75~400a short paragraph
1,000~750~4,000a long blog section
100,000~75,000~400,000a short book
1,000,000~750,000~4,000,000~10 novels
Why should anyone publishing content care? Because tokens are the unit behind almost everything an AI system does to your text: how much of your page a model can hold at once, how its owner gets billed, and why it sometimes mangles a number or a brand name. Understanding tokens is how you tell the real constraints apart from the [large language model](https://geotoolbox.ai/glossary/large-language-model) folklore. If you searched "AI tokens" and landed on coin prices, that is a different meaning entirely. Crypto "AI tokens" are tradeable assets tied to AI projects. The tokens in this article are units of text. Same word, unrelated topic. ## How Tokenization Works: From Text to Token IDs
![Hello, world! as text versus the four token IDs a model receives.](/blog/what-are-tokens-in-ai/text-to-token-ids.png)
The tokenizer turns "Hello, world!" into four integers before the model ever sees it.
Before your text ever reaches the model, a separate program called a **tokenizer** chops it into tokens and replaces each one with a number. "Hello, world!" does not enter the model as words. It enters as a short list of integer IDs, something like [9906, 11, 1917, 0] (those are GPT-4's tokenizer; a newer model assigns different numbers). The model works with the numbers. The model usually receives tokenizer IDs rather than a native character stream; the useful character-level structure is not directly exposed as separate letters. The splitting follows a method called **subword tokenization**. Frequent words get their own single token. Rare words, brand names, and long compounds get broken into pieces. "Tokenization" might split into "token" and "ization." A made-up product name might shatter into four or five fragments, which is why models sometimes garble an unusual brand name: they are rebuilding it from parts, not recalling it whole. The tokenizer is not reading for meaning, just matching against a fixed vocabulary of known chunks built once, in advance. Images and audio get the same treatment in multimodal models, sliced into patch and audio tokens, so the logic here carries over. That vocabulary comes from an algorithm called **byte pair encoding (BPE)**, a 1990s data-compression trick that Sennrich, Haddow, and Birch [adapted for neural text models in 2016](https://arxiv.org/abs/1508.07909). BPE starts from individual characters and repeatedly merges the most common neighboring pairs until it has a vocabulary of the desired size. GPT-2 settled on 50,257 tokens. The GPT family still uses BPE, through OpenAI's [tiktoken](https://github.com/openai/tiktoken) library (the `cl100k_base` vocabulary for GPT-4, `o200k_base` for GPT-4o). Other model families use close cousins: WordPiece for BERT, SentencePiece for Gemini- and earlier Llama-class models. Here is the part that matters for everything downstream. Tokenizing is only step one. Each token ID is then mapped to a [vector embedding](https://geotoolbox.ai/blog/vector-embeddings), the list of numbers that actually carries meaning. Token, then ID, then vector. The token is the raw cut. The embedding is [where understanding starts](https://geotoolbox.ai/blog/how-does-chatgpt-work). ## Why ChatGPT Can't Count the R's in "Strawberry" The famous failure where a model insists "strawberry" has two R's comes straight from tokens. The word arrives as a handful of subword tokens, and not one of them is a letter. The model usually receives tokenizer IDs rather than a native character stream; for a word like "strawberry," the useful character-level structure is not directly exposed as separate letters. It sees two or three opaque IDs and is asked to count something it cannot look at directly. So spelling questions are really memory tasks for a model, not perception. To count the R's, it has to recall how the word is spelled from its training data, then count over the recalled letters, and both steps can go wrong. It is like being asked how many times the letter E appears in a word you have only ever heard out loud. Tokens are not the whole story, to be fair. Ask a model to spell "strawberry" and it usually can, then it still miscounts, which means the counting step fails on its own. Tokens are why the task is hard. They are not the only reason it goes wrong. Arithmetic breaks for a related reason. Long numbers get chopped into tokens at arbitrary points, so "1234567" might split into pieces that do not line up for digit-by-digit math. That is part of why models have fumbled questions like whether 9.11 is bigger than 9.9. And there is only a fixed amount of computation per token, nothing like the open-ended loop you would use to carry digits through a long sum. Newer models look better at this, but be precise about why. A common workaround is to spell the word out or separate the letters, making the character-level task easier for the model. It is a workaround, not a fix at the tokenizer level. As of late 2025, [reports were still catching GPT-5.2 miscounting](https://dataconomy.com/2025/12/15/gpt-5-2-still-counts-two-rs-in-strawberry/) some variants. By mid-2026 the newest flagships usually pass the famous "strawberry" test, largely because that exact case is now familiar to them. They can stumble on less familiar, misspelled, or invented words, since subword tokenization keeps the letters out of direct reach. This is the same family of gap behind other [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations): the model confidently reports something its architecture did not actually let it check. The same blind spot explains a pain every writer has hit: ask for 1,000 words and you often get 700, with the model insisting it delivered 1,000. It generates token by token with no running word counter, so it cannot track its own length any better than it can count R's. Treat anything character-level or count-based the same way, from reversing a string to solving Wordle, and lean on models for meaning rather than spelling. ## Tokens, Context Windows, and Memory A model's **context window** is the most you can put in front of it at once, and it is measured in tokens, not words or pages. Everything has to fit inside that budget: your prompt, any documents you paste, the system instructions you never see, and the answer the model is about to write. When people say a model "remembers" a long conversation, what they mean is the whole conversation still fits in the window. These windows have grown fast, and they keep moving, so treat any single number as a snapshot. The trajectory by era:
Model (era)Approx. context window
GPT-3 (2020)~2,048 tokens
GPT-4 (2023)~8,000 to 32,000 tokens
GPT-4o (2024)~128,000 tokens
Claude 3 and 3.5 (2024)~200,000 tokens
Gemini 1.5 Pro (2024)up to ~1 to 2 million tokens
Frontier models (2026)~1 million, a few far higher
By mid-2026 a [dozen-plus frontier models ship windows of a million tokens or more](https://artificialanalysis.ai/models), and the largest open-weight model advertises 10 million. One honest caveat the marketing skips: usable context runs smaller than the advertised number, so a million-token window does not buy a million tokens of reliable attention. When you blow past the window, the system does not warn you politely. An API may throw a context-limit error, while chat products may truncate, summarize, compact, or otherwise manage older context, which is why a long session can seem to forget how it started or lose the top of a document you pasted. You cannot count on any one behavior. The only guarantee is that the full original text no longer fits. This is also where tokens meet AI search. When ChatGPT or a Google AI Overview answers a question, [retrieval-augmented generation](https://geotoolbox.ai/blog/what-is-rag) fetches passages from the web and stuffs them into that same token budget before the model writes a word. The [context window](https://geotoolbox.ai/glossary/context-window) is finite, so the system keeps only the passages it ranks highest. Your content is competing for room measured in tokens. ## How Token Pricing Works, and Why Each Reply Costs More When a company builds on an AI model through its API, it pays by the token, usually quoted as a price per million tokens. Two details surprise people. First, **output tokens cost more than input tokens**, often several times more, because generating text one token at a time is the expensive part. Reading your prompt is cheap. Writing the answer is not. Second, a chat does not bill only your latest message. To answer turn five, the model re-reads turns one through four, so the whole conversation so far rides along and is charged again, which is why a long back-and-forth gets more expensive with every reply. It is also why an agent that [reloads its whole context on every step](https://geotoolbox.ai/blog/reduce-claude-code-token-costs) can run up a bill fast. Providers now discount repeated text through prompt caching, but the structure stands: the history travels with every turn. Hidden inputs add up the same way, from the system prompt to any internal reasoning tokens a model burns before its visible answer. There is one real cost lever here, and it is worth naming precisely so it does not get misapplied. Trimming filler out of prompts and instructions can cut API spend, with [one security firm putting the savings around 10 to 30%](https://www.pivotpointsecurity.com/ai-tokens-how-they-impact-usage-costs/). That is a genuine practice, but it lives entirely on the application-building side. It is about the prompts a developer sends, not the web copy you publish. Hold that distinction, because the SEO world routinely blurs it. There is no fixed answer, since rates differ by model and change often, but the order of magnitude holds. At the few-dollars-per-million-tokens rates common in 2026, a dollar buys somewhere in the hundreds of thousands of input tokens, on the order of a few hundred pages of text. Output, priced higher, buys less. One more clarification, since it is the question behind a lot of confusion: a ChatGPT Plus subscription is a flat monthly fee, not per-token billing. Per-token pricing is the API world. Most people writing content never touch it directly. ## The Non-English Token Penalty Tokenizers are not neutral across languages, and the gap is large. The same sentence translated out of English can take far more tokens to represent, because the tokenizer's vocabulary was trained mostly on English text and has fewer ready-made chunks for everything else. A [study by Petrov and colleagues](https://arxiv.org/abs/2305.15425) found the token count for the same content can run up to roughly 15 times longer in some languages than in English. The penalty scales with how far a language sits from English and which tokenizer is doing the cutting. Many European languages land around one and a half to three times the token count. Languages written in non-Latin scripts, like Arabic and Hindi, run higher still, and severely under-resourced languages fare worst. In some cases a word produces more tokens than it has letters. For anyone publishing or budgeting AI work in more than one language, this is a real line item. The same content costs more to process, and it eats more of the context window, in nearly every language the studies have measured. ## Do Tokens Affect Your AI Search Visibility? This is where the topic gets sold badly, so here is the honest answer. **Tokenization is not a dial you can turn.** You do not pick the tokenizer. You can inspect an open one like OpenAI's to see how it splits a sample, but you cannot change it, and the closed retrieval pipelines behind AI search split and select pages in ways they never publish. Any advice that tells you to write "for the tokenizer" is selling a control that does not exist on your side. What is real sits one layer up. AI search shortlists content by meaning, comparing the embeddings of your passages against a query using [semantic search](https://geotoolbox.ai/glossary/semantic-search). Passages that are clear and self-contained tend to retrieve better. But that is just good writing. It is the same advice that worked before anyone said the word token, and it has a mechanism behind it, not a trick. A few claims to retire. "Chunk your content into token-sized pieces" is one Google's own [AI features guidance](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) now calls unnecessary, the same point we make about [content chunking](https://geotoolbox.ai/blog/content-chunking). "Token-optimize your copy to rank in AI" borrows the API cost practice from the pricing section and pretends it applies to web content, with no evidence behind it. "Pick a token-friendly brand name so AI spells it right" is another: a heavily split name can wobble, but renaming your company around a tokenizer you cannot see is not a strategy. And "one token equals one word" is simply wrong, as the first table showed. In our experience at geotoolbox, the people most confused here have read genuine engineering advice about cutting token costs in an app and assumed it must apply to their blog. It does not. What you actually control sits above the tokenizer, in whether your pages are clear, retrievable, and reachable, which is what our guide to [optimizing for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) is about. ## Frequently Asked Questions ### How many words is 1,000 tokens? About 750 words of English. The rule of thumb is one token to roughly three-quarters of a word, so a 1,000-token reply runs about two paperback pages. Punctuation, rare words, and other languages shift the count, but 750 is close enough to plan around. ### Why can't ChatGPT count the letters in "strawberry"? Because the word reaches it as a few subword tokens, with the useful character-level structure not directly exposed as separate letters, so counting R's is recall-and-tally from memory rather than something it can read off the page. Models that answer correctly usually write the word out letter by letter first, which is a workaround, not a cure. ### How do I check how many tokens my text uses? Use a tokenizer tool. OpenAI's free [Tokenizer](https://platform.openai.com/tokenizer) shows the exact split for GPT models, and its [tiktoken](https://github.com/openai/tiktoken) library does the same in code. For other model families, Hugging Face's Tokenizer Playground covers most open tokenizers. Counts differ between families, so a number from one tokenizer is only an estimate for another. ### Do I pay for tokens in ChatGPT? Not in the consumer app. A ChatGPT subscription is a flat monthly fee with usage limits. Per-token billing is the API world, where a business pays separately for the tokens it sends in and the tokens the model writes back. ### What happens when I hit the token limit? In an app you get a "context length exceeded" error; in a chat, the oldest turns usually fall away. If you hit it, start a fresh chat, paste back only the part that matters, or ask for a summary you can carry forward. A model with a larger context window buys you room, not immunity. ### Is an "AI token" a cryptocurrency? No. These are two unrelated meanings. In AI, a token is a unit of text a model processes. In crypto, "AI tokens" are tradeable digital assets attached to AI projects. This article is only about the text kind. ## What Tokens Actually Change for You Tokens are a diagnosis, not a dashboard. They explain why the machine miscounts, truncates, and bills the way it does. The controls all sit somewhere else. The earliest one is access. If an AI crawler cannot reach your page, none of this token machinery ever runs on it, because the page never enters the budget. That gate you can actually test: our free [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) shows whether the relevant search crawlers and user agents for ChatGPT Search, Perplexity, and Google Search AI features can reach a page, and what is blocking them when they cannot. Tokens tell you how the machine reads. Reachability tells you whether it reads you at all. ## Sources - OpenAI Help: What are tokens and how to count them - `help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them` - OpenAI Tokenizer (interactive tool) and tiktoken (library) - `platform.openai.com/tokenizer` - `github.com/openai/tiktoken` - Sennrich, Haddow & Birch (2016): Neural Machine Translation of Rare Words with Subword Units (byte pair encoding) - `arxiv.org/abs/1508.07909` - Petrov et al. (2023): Language Model Tokenizers Introduce Unfairness Between Languages - `arxiv.org/abs/2305.15425` - Dataconomy: GPT-5.2 still counts two R's in strawberry - `dataconomy.com/2025/12/15/gpt-5-2-still-counts-two-rs-in-strawberry` - Google Search Central: Guidance on AI features and generative AI - `developers.google.com/search/docs/fundamentals/ai-optimization-guide` - Pivot Point Security: AI tokens and how they impact usage costs - `pivotpointsecurity.com/ai-tokens-how-they-impact-usage-costs` - Artificial Analysis: AI model comparison, including context windows (2026) - `artificialanalysis.ai/models` --- ## What Is RAG (Retrieval-Augmented Generation)? > RAG (retrieval-augmented generation) is the engine behind AI search. What it is, how it works, and how to be the content that gets retrieved and cited. - Canonical: https://geotoolbox.ai/blog/what-is-rag - Published: 2026-06-11 · Updated: 2026-07-26 RAG, short for retrieval-augmented generation, is the technique that lets an AI model look things up before it answers instead of relying only on memory. It is also the machinery behind AI search. When ChatGPT, Perplexity, or a Google AI Overview answers a question and cites a few pages, retrieval-augmented generation is why those pages got pulled in. Most explainers cover RAG for the people building it. This one is for the people on the other end of it: anyone who publishes content and wants to understand why some pages get retrieved and cited while others never do. ## What Is Retrieval-Augmented Generation (RAG)? **Retrieval-augmented generation (RAG)** is a technique that lets a large language model look information up at answer time instead of relying only on what it memorized during training. The RAG system retrieves relevant documents; the model then reads the supplied context and writes an answer grounded in it. The cleanest way to picture it is an open-book exam. A plain language model takes a closed-book exam: it answers from memory, and when memory fails it guesses confidently. RAG hands the same model the textbook and lets it check the relevant page before answering. The knowledge it uses no longer has to be baked into its weights. It can be pulled from a source the moment the question is asked. That open-book step is also where your content enters the picture. When the model goes looking for a page to ground its answer, it is running a retrieval contest, and your page is either in the running or it isn't. You have almost certainly seen the output already: an AI answer with a handful of sources linked underneath is retrieval-augmented generation in action. The name describes the sequence exactly: **retrieve** the relevant documents, **augment** the prompt with them, then **generate** the answer. Keep those three words in order and the rest of RAG follows from them. ## How RAG Works: Retrieve, Augment, Generate
![Three-step RAG flow: retrieve matching passages, augment the prompt, generate a grounded answer.](/blog/what-is-rag/rag-retrieve-augment-generate.png)
The RAG loop: find the passages, paste them into the prompt, write an answer grounded in them.
A RAG system runs three steps every time someone asks a question. **Retrieve.** The system turns the user's question into a search and pulls the most relevant passages from a knowledge base. That knowledge base is usually a set of documents that have been split into chunks and converted into [vector embeddings](https://geotoolbox.ai/blog/vector-embeddings), numerical representations that let software compare meaning rather than match exact words. An embedding model converts the user query into a vector too, and the retriever finds the passages whose vectors sit closest to it. This is [semantic search](https://geotoolbox.ai/glossary/semantic-search), often combined with old-fashioned keyword matching for the terms that have to be exact. **Augment.** The retrieved passages get pasted into the prompt alongside the original question. The model now sees the user's words plus a few paragraphs of supporting evidence it did not have a second ago. Nothing about the model has changed. It just has more context in front of it for this one request. **Generate.** The model writes its answer using that supplied context, and a well-built system asks it to cite which passages it leaned on. Here is the part most explainers skip, and the part that matters most if you publish content. **Inside the RAG loop, the model does not learn your page. It reads it fresh, for that one answer, and forgets it the moment the response is done.** Public pages can still get absorbed into a model's weights during training, but that is slow, opaque, and not something you control or can update. Retrieval is. Your content gets fetched, used, and dropped, every single query, which is why being structured, current, and easy to retrieve matters more than being famous enough for the model to "know" you. The vocabulary trips people up here. A vector database is the common way to store those embeddings for fast retrieval, but it is an implementation detail, not part of the definition. RAG is the broad idea of pairing information retrieval with generation, grounding the answer in retrieved documents. The plumbing underneath can vary. ## Where RAG Came From The term comes from a 2020 paper, [Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks](https://arxiv.org/abs/2005.11401), led by Patrick Lewis with a team of machine learning researchers from Facebook AI Research (now Meta AI), University College London, and NYU. As of July 2026, Google Scholar shows [more than 24,000 citations](https://scholar.google.com/scholar?q=Retrieval-Augmented+Generation+for+Knowledge-Intensive+NLP+Tasks), and it is the reference point everyone else builds on. The paper's framing is still the most useful one. It describes combining **parametric memory**, the knowledge stored in a model's trained weights, with **non-parametric memory**, a searchable index the model can consult at inference time. In the original work that index was a dense vector representation of Wikipedia, reached through a neural retriever. Swap Wikipedia for "the live web" and you have a fair sketch of how AI search engines work today. Lewis has even [said he regrets the clunky acronym](https://blogs.nvidia.com/blog/what-is-retrieval-augmented-generation/), noting the team "would have put more thought into the name had we known our work would become so widespread." It did. ## Why RAG Exists: What It Fixes (and What It Doesn't) A standalone language model has predictable weak spots, and RAG was built to patch them. It has a **knowledge cutoff**. Training data is frozen at a point in time, so the model gets steadily more out of date until someone retrains it. RAG sidesteps this by fetching current information when the question is asked. It also has **no access to private or proprietary data**, the internal documents and recent pages that were never in its training set. RAG connects that external knowledge in without retraining. And because retrieval is cheap compared to fine-tuning a model on new data, it is the cost-effective way to keep answers current. The headline benefit is [grounding](https://geotoolbox.ai/glossary/grounding). By anchoring answers in retrieved sources, RAG **reduces hallucinations**, the confident, made-up answers models produce when they are working from memory alone. It also makes answers checkable, because the system can cite the passages it used. Now the honest part, because this is the single most oversold claim in the category: **RAG reduces hallucinations, it does not eliminate them.** The model can still misread a correct source, stitch together conflicting passages, or write something unsupported when retrieval comes back thin. Retrieval quality sets a hard ceiling on the whole system: a model cannot ground an answer in a passage the retriever never found. The evidence is blunt. A Stanford study, [Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools](https://reglab.stanford.edu/publications/hallucination-free-assessing-the-reliability-of-leading-ai-legal-research-tools/), tested commercial legal-research tools that are themselves RAG systems built on curated, authoritative law libraries. It found Lexis+ AI hallucinated on more than 17% of queries and Westlaw's tool on roughly a third, despite marketing that implied none at all. If purpose-built RAG over a clean legal corpus still misses that often, treat any "hallucination-free" promise with suspicion. RAG is a strong mitigation, and that is exactly where its value sits. For more on why this happens, see this breakdown of [AI hallucinations](https://geotoolbox.ai/blog/ai-hallucinations). ## RAG vs Fine-Tuning: The "Train ChatGPT on Our Docs" Confusion When a stakeholder says "let's train ChatGPT on our website," they almost always mean RAG, not training. The two get blended constantly, and picking the wrong one is expensive. **Fine-tuning** changes the model itself. You run additional training so new patterns get written into its weights. After that the knowledge is internal, there is no lookup step, and updating it means training again. Fine-tuning is the right tool for teaching a model a style, a format, or a behavior. **RAG** leaves the model untouched and gives it documents to read at answer time. Knowledge lives in a separate index you can update whenever you want, and the model cites what it pulled. RAG is the right tool for facts that change or that the model was never trained on.
Question RAG Fine-tuning
What changes? An external index of documents The model's own weights
When is knowledge added? At answer time, per query During a training run, up front
Best for Fresh or proprietary facts Style, tone, format, behavior
Updating it Edit the index, no retraining Retrain or re-tune the model
Can it cite sources? Yes Not by itself / not reliably grounded
They are not rivals. Production systems often fine-tune for behavior and use RAG for current facts. The original authors even described their method as "a general-purpose fine-tuning recipe" for building RAG models. The practical rule: if the problem is "the model does not know this fact," reach for RAG. If the problem is "the model does not answer in the way we need," reach for fine-tuning. For publishers, the distinction matters for one reason in particular: the AI engines that might cite you are running RAG, not fine-tuning on your site. Which leads to the part that actually affects your traffic. ## RAG Is How AI Search Actually Works The RAG systems most people actually interact with are the AI search engines you are already trying to show up in, and that is the connection the vendor explainers leave out. They describe RAG as enterprise plumbing: a support chatbot answering over a company's internal data, internal knowledge search, a research assistant reading proprietary files. Those use cases are real, but they are the smaller story. When ChatGPT browses the web to answer a question, it retrieves web information through ChatGPT search, which OpenAI says uses third-party search providers, partner-provided content, and its search crawling systems. [Perplexity](https://geotoolbox.ai/blog/perplexity-seo) typically retrieves and cites sources. Google's AI Overviews draw on Google's Search index and ranking/quality systems, then synthesize selected sources. Different engines, same three steps: retrieve, augment, generate. This is just [how AI search works](https://geotoolbox.ai/blog/how-does-ai-search-work) under the hood. ChatGPT itself shows how the pieces fit. The base model is a language model, but its search and browsing mode wraps that model in a RAG loop, retrieving live pages before it answers. The model is the generator; the product around it is the RAG system. That reframing changes what "getting cited" means. **If AI search is RAG, then getting cited starts with getting retrieved by a RAG pipeline.** Retrieval is the qualifying round: the model still chooses which of the retrieved sources to cite, but a page that is never retrieved cannot be cited at all. There is no list of ten blue links to scroll, just one synthesized answer with a few sources. Retrieval is the contest most pages were never written to enter. ## How to Be the Page That Gets Retrieved If you want your pages pulled into AI answers, you have to make them easy to retrieve. This is where RAG stops being trivia and turns into a content strategy, and a few rules follow directly from how the pipeline works. **Passages get retrieved, not pages.** A RAG system splits documents into chunks and retrieves the chunk that best matches the query, not your whole article. So each section has to make sense on its own. Put the answer to a question directly under a clear, question-style heading, in the first sentence or two, before the context and caveats. In our experience auditing pages for AI visibility, this is the most common fixable problem: the answer exists, but it is buried three sentences into a paragraph, and the chunk that gets retrieved is the lead-in, not the payload. Write each section so a reader who lands on it cold still gets the answer. That is the same instinct behind good [content chunking](https://geotoolbox.ai/blog/content-chunking), with one caveat: this means writing self-contained sections for readers, not chopping pages into artificial fragments for machines, a distinction we come back to below. **You cannot be retrieved if you cannot be reached.** Retrieval runs over an index, and you only enter the index if the engine's crawler is allowed to fetch you. Check that you are not blocking [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) you actually want citing you. Reachability is the floor; everything else is wasted if the page never gets fetched. **Write with unambiguous clarity.** Models misread sources that are vague or rely on outside context. State claims plainly and self-containedly so a retrieved snippet cannot be misinterpreted. **Freshness helps, with a caveat.** Recency is one of the few signals that correlates with getting cited, since the entire point of retrieval is to beat a model's stale memory. Genuine updates, not date-bumping, are worth making. Just treat freshness as a correlation, not a guaranteed lever.
What to do Why it helps retrieval
Answer-first sections under clear headings The retrieved chunk contains the answer, not the wind-up
Self-contained paragraphs (one idea each) A chunk still makes sense pulled out of context
Allow the AI crawlers you want citing you You can only be retrieved if you are in the index
Plain, specific, unambiguous claims Lowers the chance a passage gets misread
Keep pages genuinely current Recency correlates with being pulled into answers
Ranking well and getting cited are not the same job. Ranking can get a page considered, but retrieval decides whether a passage of yours actually gets quoted, so a page can sit at the top of Google and still never appear in an AI answer. And no, this is not just SEO with a new logo, though it is closer than the hype suggests. Google itself says its systems [can read long, multi-topic pages and extract the relevant passage](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) without you chopping content into artificial fragments, and that the fundamentals of [GEO and AEO are still SEO](https://geotoolbox.ai/blog/geo-vs-aeo-vs-seo). The honest summary: the same fundamentals (topical depth, authority, clarity), plus a real shift in the unit of retrieval from the page to the passage. If you want the full playbook, we cover it in [how to optimize for AI search](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) and [writing pages LLMs cite](https://geotoolbox.ai/blog/ai-content-optimization). ## Is RAG Dead? Agentic RAG and the Long-Context Debate Neither huge context windows nor the shift toward AI agents kills retrieval, whatever the "RAG is dead" headlines say. Those are the two reasons usually given: models that can now swallow a whole document set in a single prompt, and more autonomous agents doing the work. The slogan is really about **naive** RAG, the simplest version that retrieves once and generates once. That basic pipeline is often not enough for multi-step work, and the industry is moving toward **agentic RAG**, where an agent decides when to retrieve, reformulates the query, retrieves again, and checks what it got back before answering. That is more retrieval with better judgment around it, not less. Long-context models are a genuine alternative for some jobs, but pasting everything into the prompt is slower and more expensive than retrieving the few passages that matter, and it does not scale to the open web. So retrieval stays central. For anyone who publishes content, the practical takeaway does not change. AI search engines still retrieve before they answer. Whether the system is naive or agentic, it has to find your page to cite your page. Being retrievable is still the price of admission. ## Frequently Asked Questions ### Is ChatGPT a RAG model? The underlying model is not, but ChatGPT's search and browsing mode is. When ChatGPT looks something up before answering, it retrieves live web pages, adds them to the prompt, and generates a grounded reply. That retrieve-augment-generate loop is RAG, with the model acting as the generator inside it. ### Does RAG stop hallucinations? It reduces them, it does not stop them. Grounding answers in retrieved sources lowers the rate of made-up content, but the model can still misread a source or fill gaps when retrieval comes back weak. A Stanford study found commercial legal RAG tools still hallucinated on 17% to 33% of queries, so treat any "hallucination-free" claim with caution. ### What is the difference between RAG and a vector database? RAG is the overall technique of grounding a model's answer in retrieved documents. A vector database is one component some RAG systems use to store embeddings and run fast similarity search. You can build RAG without one, using keyword search or a knowledge graph instead, so the vector database is a common piece of plumbing, not a requirement. ### What are the types or "levels" of RAG? People usually describe a progression: naive RAG (retrieve once, generate once), advanced RAG (better chunking, reranking, and hybrid search to improve what gets retrieved), and agentic RAG (an agent decides when and what to retrieve and can iterate). They are points on a spectrum, not rigid categories. A separate, complementary idea skips retrieval for a small set of curated facts and hands the agent a structured knowledge file directly, which is the pitch behind [OKF vs RAG](https://geotoolbox.ai/blog/okf-vs-rag). ### Do I need to build a RAG system to benefit from AI search? No. Most publishers are on the receiving end of someone else's RAG, not building their own. Your job is to make your existing content easy to retrieve and cite, not to stand up a pipeline. ## You Don't Get Into the Model. You Get Retrieved. So the practical work is short, and it is ongoing: keep pages reachable, write sections that stand alone as clean answers, stay current, and give each one a direct answer worth quoting. Do that and you are optimizing for the retrieval step every AI engine runs, instead of chasing a ranking that may never turn into a citation. The first thing to check is whether AI engines can even reach and read your pages, because nothing else matters if they can't. That is exactly what our [AI Readiness check](https://geotoolbox.ai/tools/ai-readiness) looks at, and from there you can [measure how often you're actually getting cited](https://geotoolbox.ai/blog/how-to-track-ai-visibility) across the engines that run on retrieval. ## Sources - Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks - Lewis et al., 2020 (NeurIPS 2020) - `arxiv.org/abs/2005.11401` - What Is Retrieval-Augmented Generation, aka RAG? - NVIDIA, 2023 (updated 2025) - `blogs.nvidia.com/blog/what-is-retrieval-augmented-generation` - Retrieval-augmented generation - Wikipedia - `en.wikipedia.org/wiki/Retrieval-augmented_generation` - Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools - Magesh et al., Stanford RegLab & HAI, 2024 - `reglab.stanford.edu/publications/hallucination-free-assessing-the-reliability-of-leading-ai-legal-research-tools` - Guide to Optimizing for Generative AI Features on Google Search - Google Search Central - `developers.google.com/search/docs/fundamentals/ai-optimization-guide` --- ## Content Chunking: What It Means and What Google Says to Skip > Content chunking is a real RAG-pipeline concept and a shaky SEO tactic. What Google's AI guide says to skip, and what page structure still earns citations. - Canonical: https://geotoolbox.ai/blog/content-chunking - Published: 2026-06-10 · Updated: 2026-07-26 Content chunking is having a strange year. A wave of SEO advice sells it as the key to AI visibility. Google's official guidance now lists it among the tactics to ignore. Both camps use the same word for different things, which is how a real engineering concept became a shaky content tactic. This page sorts it out: what chunking actually is inside AI retrieval, what Google said and meant, which claims survive an evidence check, and what to do with your pages instead.
![Three meanings of content chunking compared: RAG pipeline step, UX principle, and the SEO hack.](/blog/content-chunking/content-chunking-three-meanings.png)
Same word, three jobs — and only one of them is something a page author actually controls.
## What Content Chunking Actually Means (Three Different Things) **Content chunking** means three different things depending on who is talking, and most of the bad advice comes from mixing them up. In AI retrieval engineering, chunking is a pipeline step: a [retrieval-augmented generation (RAG)](https://geotoolbox.ai/blog/what-is-rag) system splits documents into segments, converts each segment into an embedding, and pulls back the most relevant segments when a query comes in. The system does the splitting. The document's author is not in the room. In UX and psychology, chunking is about human working memory. [Nielsen Norman Group defines it](https://www.nngroup.com/articles/chunking/) as breaking content into small, distinct units of information so people can process it, building on George Miller's 1956 finding that most people hold roughly seven chunks in short-term memory. This sense owns most of the Google results for the term, and it has nothing to do with AI search. Then there is the third sense: content chunking for AEO and GEO purposes, the tactic version. "Chunk your content into AI-sized pieces and engines will pick you." This version borrows the vocabulary of the first sense and the credibility of the second, and it is the one Google now explicitly tells publishers to ignore.
SenseWho uses itWhat it meansWho controls it
RAG pipeline stepAI engineersSplitting documents into segments for embedding and retrievalThe engine, not you
UX writing principleDesigners, instructional writersBreaking information into digestible units for human memoryYou
AEO/GEO hackParts of the SEO industryPre-fragmenting pages so AI systems "extract" youNobody, which is the problem
This article is about the first and third senses: what chunking really does inside AI search, and why the hack version does not survive contact with the evidence. ## How Chunking Really Works Inside AI Retrieval When an AI search engine processes your page, the chunking happens on its side of the fence. This engine-side splitting is what chunking in AI originally means: the pipeline fetches your content, splits it into chunks, converts each chunk into [vector embeddings](https://geotoolbox.ai/blog/vector-embeddings), and stores them in an index. At query time, the system retrieves the chunks that best match the question, often after [fanning the query out](https://geotoolbox.ai/blog/query-fan-out) into several sub-queries. We cover the full retrieval flow in our breakdown of [how AI search retrieval works](https://geotoolbox.ai/blog/how-does-ai-search-work). ### The Engineer's Dials, Not Yours The part that matters for this debate: chunking strategy is an engineering setting, not a property of your page. [Amazon Bedrock's documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/kb-chunking.html) shows what this looks like in practice. A developer building a knowledge base picks fixed-size chunking, semantic chunking, hierarchical chunking, or no chunking at all, and tunes token counts and overlap percentages. Bedrock's default splits text into chunks of roughly 300 tokens while honoring sentence boundaries. Those are the engineer's dials. You never touch them. Research on RAG chunking treats it the same way: as a system trade-off for the pipeline owner. A [systematic analysis of chunking strategies](https://arxiv.org/abs/2601.14123) (Bennani and Moslonka, preprint, January 2026) varied chunking method, chunk size, and overlap across a standard RAG setup and found that chunk overlap provided no measurable benefit while increasing indexing cost. The interesting part is the framing. The paper's audience is engineers deploying retrieval systems. Nothing in it suggests document authors should write differently. ### There Is No Universal Chunk Here's the nuance worth keeping. Engineers choose among many splitting strategies, including fixed-size, sliding-window, recursive, semantic, and late chunking, and some of them do read document structure: structure-aware and HTML-aware chunking split at headings and sections, and semantic chunking looks for topic shifts. So a well-organized page can interact better with some pipelines. But which strategy any given engine uses, at what size, with what overlap, is invisible to you and different across engines. There is no universal chunk, so there is nothing stable to optimize against. ## The Chunking Hack Google Just Told You to Ignore Google has now said it directly. Its [guide to optimizing for generative AI features](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide), last updated July 2026, lists "'Chunking' content" among the tactics publishers can skip: "There's no requirement to break your content into tiny pieces for AI to better understand it. Google systems are able to understand the nuance of multiple topics on a page and show the relevant piece to users." If your pages run long and cover several topics, that is not a defect to fix. Chunking sits in that ignore list next to [llms.txt files](https://geotoolbox.ai/blog/llms-txt), AI-specific rewriting, and overfocusing on structured data. The same guide we examined when testing [schema markup for AI search](https://geotoolbox.ai/blog/schema-markup-for-ai) applies here: Google says tactics like chunking content, unnecessary AI text files, and special generative-AI schema are not required; ordinary SEO foundations still matter. This was not a one-off documentation edit. On the Search Off the Record podcast in January 2026, Google's Danny Sullivan addressed the bite-sized-chunks advice [in unusually plain terms](https://searchengineland.com/google-doesnt-want-you-to-create-bite-sized-chunks-of-your-content-467269): "So we don't want you to do that. I was talking to some engineers about that. We don't want you to do that. We really don't." Per Search Engine Land's report, he added that Google does not want publishers producing one version of their content for the LLM and another for the web, and that ranking systems keep improving toward rewarding content written for humans. So where did the tactic come from? The likeliest origin story runs through RAG tutorials. As chunking became a visible concept in AI engineering, its vocabulary leaked into SEO advice, where "the pipeline splits documents into chunks" mutated into "you should split your page into chunks." Productized restructuring retainers did the rest, and a pipeline internal became a billable deliverable. To be clear about where we stand, since geotoolbox sells AI-visibility measurement: auditing how AI systems actually treat your pages is real work. Restructuring pages to hit chunk specs no engine publishes is the part we would call a tax on confusion. ## The Chunk-Size Numbers Nobody Can Tie to AI Chunking advice comes wrapped in oddly specific numbers. From [one widely shared SEO guide](https://searchengineland.com/guide/content-chunking-seo) and its peers: - Keep passages to 100-300 words - Aim for 500-token chunks - Build macro chunks of 300-800 words and micro chunks of 100-200 - Place a heading every 200-300 words - Engines extract 40-60 word excerpts Try to trace any of these to a study of AI retrieval and you come up short. I found no published evidence tying these specific numbers to AI citation gains. What you find instead is recycling: real numbers from the wrong context. The token counts are defaults from RAG configuration docs, settings for a system the publisher does not run. The 40-60 word figure descends from featured-snippet research that measured what Google displays in a snippet box, not what AI systems prefer to retrieve. The heading-frequency rule matches the readability threshold a popular SEO plugin has flagged for years. Each number is real somewhere. None is a measured property of AI retrieval, and quoting them as AI-citation targets is like telling drivers to inflate their tires to the PSI spec of someone else's truck. Even the one well-documented benchmark in this space points the other way. [NVIDIA tested chunking strategies](https://developer.nvidia.com/blog/finding-the-best-chunking-strategy-for-accurate-ai-responses/) across five test datasets and found page-level chunking, treating the whole page as the retrieval unit, delivered the highest average answer accuracy (0.648) of everything tested. That is a benchmark for engineers building internal RAG systems; web publishers are not its audience, and NVIDIA's own refinements vary by corpus and query type. But notice what it implies: in the best-documented test available, bigger chunks won. The micro-fragmentation advice points in the opposite direction from its own favorite evidence genre. And beneath all of it sits [tokenization](https://geotoolbox.ai/blog/what-are-tokens-in-ai). [Mark Williams-Cook's invalid-schema experiment](https://markwilliamscook.substack.com/p/schema-llms-and-the-low-bar-for-evidence) showed AI assistants happily reading deliberately broken JSON-LD because LLM assistants may still use visible text from broken JSON-LD, which is a weak basis for claiming schema directly drives AI citations. That does not mean search systems never parse structured data. His broader point is the one this whole debate needs: "Be ruthless about the bar of evidence." In our experience tracking which pages get cited across AI engines at geotoolbox (an observation from our tracking data, not a controlled study), the winners are not the pages hitting word-count targets. They are the ones where a retrieved passage happens to contain a complete, specific, sourced answer. The numbers game optimizes the wrong layer. ## What the Evidence Actually Supports The honest reading is not "chunking is fake." Retrieval is real, and AI answers are often grounded in specific information drawn from retrieved pages, sections, or passages. The question is which claims about that hold up, so here is the chunking debate sorted by evidence tier.
ClaimVerdictEvidence
AI systems may retrieve pages, sections, or passages, then use specific information from themConfirmed mechanismDocumented across RAG systems; Google's own guide describes reviewing "specific information from those retrieved pages"
Q&A and structured formats match queries better than dense proseTested, small scaleChris Green's embedding test: Q&A format had the highest semantic match in every scenario, dense prose the worst. His own caveat: small controlled test, and vector similarity is not ranking
Self-contained sections survive extraction betterConfirmed mechanismA retrieved passage is read without its surroundings; a passage that depends on "as mentioned above" arrives broken. This is retrieval logic, not a study result
Specific chunk-size and heading-frequency targetsClaimed, unsourcedNo study connects any circulating number to AI citations; each traces to a non-AI context (RAG config defaults, featured-snippet display stats, readability-tool thresholds)
You can pre-chunk your page to control engine cut pointsRejectedAhrefs' analysis: chunking happens inside model pipelines, every engine slices differently, and you cannot know which strategy any engine runs; in a fixed-size pipeline your section can land mid-chunk or split across two no matter how it is formatted
Good chunking trains AI or prevents hallucinations about your brandRejectedConflates retrieval-time splitting with model training; no mechanism connects your fragment sizes to what a model learns in training
Notice that the two "confirmed" rows describe the engine's behavior and the writing's properties, while everything rejected describes attempts to control the engine. That is the line through this whole topic. It also resolves the apparent contradiction that confuses people: Google saying "don't chunk" and AI systems demonstrably retrieving passages are both true. The engine chunks. You cannot chunk for it. What you can do is write passages that hold up no matter where the knife falls, which is the core habit behind [semantic SEO](https://geotoolbox.ai/blog/semantic-seo). One more disambiguation, because it trips up even careful readers: enterprise teams really do tune chunking for their own internal knowledge bases, and vendors publish guidance for that. That advice is for people who own the pipeline. The public web is somebody else's pipeline, every engine's settings differ, and none of them publish the dials. For the internal-knowledge case specifically, there is now a structured alternative to chunking curated facts at all: Google's Open Knowledge Format hands an agent whole concepts instead of retrieved fragments, which we compare in [OKF vs RAG](https://geotoolbox.ai/blog/okf-vs-rag). ## What Survives: Structure That Helps Without the Hack Strip away the chunk-size folklore and a short list of structural habits remains standing: - Answer-first openings, where the first sentence under a heading answers the heading - Self-contained sections that make sense lifted out of the page - Headings that describe what the section actually says - One idea per passage - Tables and numbered lists where the facts are genuinely list-shaped, because comparisons and sequences survive extraction better as structure than as prose Here is the difference in one before-and-after. Before: "As we covered earlier, the same problem applies to longer documents." After: "Fixed-size chunking also fails for long contracts, because clause boundaries rarely align with token windows." The first sentence dies when a pipeline lifts it away from the page; the second survives anywhere it lands. These habits survive for a reason worth being precise about: they are properties of the writing, not settings aimed at any engine's chunker. A self-contained passage works whether the pipeline slices at 300 tokens or takes the whole page, because there is no cut point that leaves the reader, human or model, holding a fragment that depends on text it cannot see. Which is also why they were good advice before AI search existed. Editors have pushed answer-first structure for decades because readers scan. The result? You do not need a chunking workflow. You need the writing craft, and we have already covered it in depth: our guide to [writing passages LLMs cite](https://geotoolbox.ai/blog/ai-content-optimization) handles the sentence-level work, and the [answer-first restructuring playbook](https://geotoolbox.ai/blog/how-to-optimize-for-ai-search) walks through retrofitting existing pages section by section. The honest caveat: nobody has published before-and-after data on restructuring projects, so anyone who claims a guaranteed citation lift from reformatting is ahead of the evidence. ## Measure Instead of Guessing The chunking debate is really a proxy for a habit: adopting tactics whose effect you cannot observe. You will never see where an engine cut your page. You can see whether engines cite it. Track citations on a set of pages before and after a rewrite, against pages you left alone. Keep the caveat in view: this is observational data and engines change underneath you, so treat movement as signal rather than proof. It still beats optimizing a layer you cannot observe at all. And being cited without earning clicks is its own measurement problem, one we cover in our guide to [tracking AI visibility](https://geotoolbox.ai/blog/how-to-track-ai-visibility). That is the check worth running before any restructuring project. geotoolbox's [Content Analyzer](https://geotoolbox.ai/features/content-analyzer) gives your pages a Citability and an AI Readability grade, built on the signals that survive this article's evidence table: answer-first capsules, self-contained sections, sourced data, real tables. Fix the pages with a measurable defect instead of reformatting everything to folklore specs. If the chunking argument lands in your team's Slack again, an audit of your own cited and uncited pages ends it faster than another opinion. ## FAQ ### Does Google's advice apply to ChatGPT and Perplexity too? Google's guide speaks for Google, including AI Overviews and AI Mode; no other engine has published equivalent guidance. The mechanism argument still generalizes: ChatGPT and Perplexity run their own retrieval pipelines with their own unpublished chunking settings, so there is still no spec to write toward. What differs between engines is which pages they retrieve and cite, not your ability to control their cut points. ### Does Google penalize chunked content? No. Google's guide calls chunking unnecessary, not harmful, and short sections are not a ranking problem. The damage shows up elsewhere: pages diced into fragments to satisfy an imagined extractor read worse for humans, and Danny Sullivan's warning was aimed exactly at publishing machine-bait instead of writing for people. ### Should I rewrite my existing pages to fix their chunking? Not as a chunking project. No before-and-after data exists on re-slicing pages; the closest published evidence, [the GEO benchmark study](https://arxiv.org/abs/2311.09735), tested adding material like statistics, quotations, and citations rather than reformatting what was already there. Start from measurement: protect pages that already earn AI citations, then fix pages where the answer to the title question is genuinely buried. A buried answer is a real defect; a 400-word section is not. ### What is semantic chunking? A pipeline technique that splits documents at topic boundaries detected with embeddings, instead of at fixed token counts. Amazon Bedrock offers it as a configurable option, with buffer sizes and breakpoint thresholds set by whoever builds the knowledge base. Like every chunking strategy, it runs engine-side; you cannot opt your pages into it. ### What is context chunking? A family of retrieval-pipeline techniques that keep chunks from arriving orphaned: semantic chunkers can embed each sentence with a buffer of its neighbors, and hierarchical setups retrieve a small child chunk but hand the model its larger parent section. Amazon Bedrock documents both. Either way, it is configured by pipeline builders; a writer never touches it. ## Sources - Google: Optimizing your website for generative AI features on Google Search - `developers.google.com/search/docs/fundamentals/ai-optimization-guide` - Search Engine Land: Google doesn't want you to create bite-sized chunks of your content - `searchengineland.com/google-doesnt-want-you-to-create-bite-sized-chunks-of-your-content-467269` - Ahrefs: SEO "Chunk Optimization" is Overrated - `ahrefs.com/blog/seo-chunk-optimization` - Chris Green: How Content Structure Matters for AI Search - `chris-green.net/post/content-structure-for-ai-search` - NVIDIA: Finding the Best Chunking Strategy for Accurate AI Responses - `developer.nvidia.com/blog/finding-the-best-chunking-strategy-for-accurate-ai-responses` - AWS Bedrock documentation: How content chunking works for knowledge bases - `docs.aws.amazon.com/bedrock/latest/userguide/kb-chunking.html` - Mark Williams-Cook: Schema, LLMs and the Low Bar for Evidence in GEO - `markwilliamscook.substack.com/p/schema-llms-and-the-low-bar-for-evidence` - Bennani & Moslonka: A Systematic Analysis of Chunking Strategies for Reliable Question Answering (arXiv preprint) - `arxiv.org/abs/2601.14123` - Nielsen Norman Group: How Chunking Helps Content Processing - `nngroup.com/articles/chunking` --- ## Schema Markup for AI Search: What It Does and Doesn't Do > Schema markup for AI search, minus the hype: which engines actually use JSON-LD, what the studies show, the types worth adding, and how to verify it arrived. - Canonical: https://geotoolbox.ai/blog/schema-markup-for-ai - Published: 2026-06-10 · Updated: 2026-07-22 Schema markup for AI search is the most over-promised tactic in generative engine optimization right now. Vendors pitch it as a citation multiplier; skeptics call it dead code that models never read. Both are wrong in instructive ways. Two engines are on record that structured data helps them understand content, the only intervention study so far found that adding it barely moved citations, and an implementation detail as small as how you inject the markup decides which AI crawlers ever see it. This guide sorts the confirmed from the claimed, and shows how to verify your markup even arrives. ## What Schema Markup Actually Does in AI Search Schema markup is JSON-LD code that labels what your page contains: this is an article, written by this person, who works for this organization, about this topic. Search engines have read it for years to power rich results. The new question is whether AI search reads it too, and the honest answer is: some engines do, some almost certainly do not, and nobody's markup turns a weak page into a cited one. The most useful way to think about it comes down to a split. The quality and authority of your content decide whether you are **worth citing**. [Structured data](https://geotoolbox.ai/glossary/schema-markup) decides, at most, how **easy you are to cite**: it removes guesswork about who you are, what the page covers, and which facts belong to which entity. That second job is real but narrow. AI systems that retrieve your page have to work out whether "Mercury" means the planet, the element, or your agency's brand name. Markup that ties the page to a defined organization, author, and topic shortens that inference. It is plumbing for machine understanding: the same job it has always done in classic SEO, carried over into [generative engine optimization](https://geotoolbox.ai/blog/what-is-geo). What schema does not do is act as an AI ranking lever. No engine has ever said "add markup, get cited," and the newest data points the other way. The citation-hack version of the pitch is the one part of this the evidence does not support. ## Does AI Use Schema Markup? Engine by Engine Two platforms have said yes on the record for their AI surfaces broadly, and one more is on record for shopping only. For citations everywhere else, silence is not a yes.
EngineUses schema?The evidence
Google AI Overviews / AI ModeReads it, doesn't require itGoogle's generative AI optimization guide (updated June 2026): structured data "isn't required for generative AI search," but keep using it for rich-result eligibility and the features that depend on it.
Microsoft Bing CopilotConfirmedFabrice Canel, principal product manager at Bing, said on stage at SMX Munich (March 2025) that schema markup helps Microsoft's LLMs understand content.
ChatGPTCitations: no confirmation. Shopping: yesOpenAI has never said its search citations use schema, but its shopping surface documents structured product metadata and a merchant product feed spec. ChatGPT search uses third-party providers and partner content; OpenAI does not publicly name Bing here. High overlap with Bing's index is a dated inference, not a platform confirmation.
GeminiNo statementGoogle's structured-data statements cover Search's AI features, not the Gemini app. Nothing published either way.
PerplexityNo confirmationNo public statement on schema anywhere in its answer pipeline.
ClaudeNo confirmationAnthropic has published nothing on structured data in its search or citation pipeline.
One more official line worth keeping verbatim, because it kills a whole category of snake oil. Per the same Google guide: "there's no special schema.org markup you need to add" for AI features, and Google's [AI features documentation](https://developers.google.com/search/docs/appearance/ai-features) adds that no AI-specific files or markup exist at all. Eligibility for an AI Overview citation is the same as for a snippet: be indexed, be crawlable. If a vendor pitches you an "AI schema type," the markup does not exist. The same guide's ignore list also covers ["chunking" content](https://geotoolbox.ai/blog/content-chunking), a sibling myth we examine separately. So the platform answer to "does AI use schema markup" is: Google and Bing, yes, as an understanding signal. Everyone else, unknown, and unknowable until they say so or someone tests it properly. Which someone now has. ## What the Evidence Says (and How Strong It Is)
![Five claims about schema and AI citations sorted from confirmed to unsupported.](/blog/schema-markup-for-ai/schema-evidence-sorted.png)
The schema-for-AI claims sorted by evidence strength — the picture is narrower than the pitch.
Claims about schema and AI visibility come in wildly different strengths, and most articles flatten them into one pile. Here is the same pile, sorted.
ClaimVerdictBasis
Google and Bing use structured data to understand contentConfirmedOfficial documentation and an on-stage statement from Bing's Fabrice Canel
LLMs can extract text into user-defined structured schemasDemonstrated in the labA Nature Communications study (February 2024) showed fine-tuned LLMs reliably extracting scientific text into user-defined schemas. Scope: trained models in a lab, not web markup, not visibility.
Pages cited by AI tend to have schemaCorrelative onlyAcross 6 million URLs, Ahrefs found AI-cited pages were almost 3x more likely to carry JSON-LD, which is where the "3x" stat in vendor decks comes from. Sites that add schema also invest in everything else; the correlation is real, the causation is what the test below checks.
Adding schema increases AI citationsNo lift in the only intervention testAhrefs tracked 1,885 pages that added JSON-LD against roughly 4,000 matched controls: Google AI Overviews -4.6%, AI Mode +2.4%, ChatGPT +2.2%, everything within noise except the small AIO decline. The study's own scope caveats: pages already cited 100+ times, 30-day windows, all schema types pooled.
"Schema triples your citations" / "40% higher CTR"Claimed, unsupportedThese numbers circulate through vendor decks and secondhand citation chains. None traces to a controlled study. Treat any precise multiplier with suspicion.
The Ahrefs result deserves careful reading rather than a victory lap in either direction. It is the only intervention test we know of, and it found nothing for pages that were already getting cited. What it cannot rule out, by its own design, is a longer-horizon or entity-establishment effect for unknown brands, which is the population schema advocates care most about. The fair summary: the burden of proof now sits with anyone claiming a citation lift. There is also a technical riddle the hype skips over: do the models even see your markup? The one public test we know of says yes, but not in the way schema fans hope. Mark Williams-Cook [planted a fake company's address](https://markwilliamscook.substack.com/p/schema-llms-and-the-low-bar-for-evidence) only inside deliberately invalid JSON-LD, and both ChatGPT and Perplexity served it back. The models had tokenized the script block as plain page text: no parsing, no validation, no credit for structure. So "tokenization strips your markup" is false, and so is "the AI parses your graph." In Williams-Cook's test, the text inside invalid JSON-LD reached ChatGPT and Perplexity through a page fetch; that test doesn't prove all systems ignore schema structure, but it does show the text pathway exists. Where structure plausibly pays is **upstream**, in the Google and Bing pipelines that are confirmed to parse it. Which raises the question almost nobody asks: does your markup survive the crawl at all? ## Does Your Schema Even Reach the Model? The Crawler Layer Before any debate about whether an engine uses your markup, there is a blunter question: did the engine's crawler receive it? This is where schema implementations quietly fail, and no validator will tell you. Here is the mechanical reality. Many [AI crawlers](https://geotoolbox.ai/blog/ai-crawlers) are raw-HTML-first or have been observed missing client-rendered content (e.g., the Vercel/MERJ crawler study); unless you have bot-specific evidence, treat JavaScript-injected schema as at-risk and verify with the exact crawler/user agent. That cuts in two directions: 1. **JSON-LD hard-coded in your HTML travels fine.** A script tag in the source is just text in the document those bots download. Whether their pipelines parse it is the open question from the table above, but at least it arrived. 2. **JSON-LD injected by JavaScript may not arrive for crawlers that do not render the page.** If your markup is added by Google Tag Manager, a consent management platform (CMP), or any client-side script, the bots that skip rendering download a page with no schema in it at all. Google can process JavaScript when it is not blocked, and Bing can process some JavaScript, but server-side JSON-LD is still safer, and that is exactly why this failure hides: the Rich Results Test shows green while the entire unconfirmed column above receives nothing. Per the Williams-Cook test, text reaching the model is the one pathway proven to work at answer time, and it is the pathway tag-manager injection forfeits. In our experience analyzing pages for AI visibility, this is one of the most common surprises we find: markup that validates perfectly in Google's tools but is absent from the HTML an AI crawler actually receives, because it ships through a tag manager. The 60-second check: open your page with view-source (not inspect, which shows the rendered DOM) or run `curl` on the URL, and search for `application/ld+json`. If your markup is only in the rendered version, move it server-side. The same JavaScript dependency that hides schema usually hides content too, which is the deeper problem covered in our guide to [agent-ready websites](https://geotoolbox.ai/blog/agent-ready-website). Schema is irrelevant if the bots cannot fetch the page at all. The free [AI Crawler Checker](https://geotoolbox.ai/tools/ai-crawler-checker) shows which of 34 AI crawlers your robots.txt allows or blocks, with the exact line to fix. ## The Schema Types Worth Implementing for AI Schema.org defines hundreds of types. For AI search, a short stack covers nearly all the value, and everything past it is diminishing returns. ### Organization, Sitewide This is your brand's identity record: name, logo, URL, and most importantly `sameAs` links to your LinkedIn, Wikipedia, Crunchbase, and other verified profiles. By the rubric above, `sameAs` is a mechanism argument rather than an engine confirmation, but it is likely the highest-value field here: it lets a system collapse scattered mentions of your brand into one entity instead of several uncertain ones. A minimal version: ```json { "@context": "https://schema.org", "@type": "Organization", "@id": "https://example.com/#organization", "name": "Example Co", "url": "https://example.com", "logo": "https://example.com/logo.png", "sameAs": [ "https://www.linkedin.com/company/example-co", "https://x.com/exampleco" ] } ``` ### Article Plus Person, on Every Post Use Article (or BlogPosting), with headline, author, dates published and modified, and publisher. This ties content to a named human and a brand, the attribution layer engines lean on for trust signals like [E-E-A-T](https://geotoolbox.ai/glossary/e-e-a-t). We ship exactly this on the page you are reading, plus BreadcrumbList and FAQPage, so the code examples here are the markup we run ourselves. ### FAQPage, with Honest Expectations Google [restricted FAQ rich results](https://developers.google.com/search/blog/2023/08/howto-faq-changes) to government and health sites in August 2023, then [retired the feature entirely in May 2026](https://developers.google.com/search/docs/appearance/structured-data/faqpage), so anyone promising FAQ dropdowns in the SERP is years out of date. Google's only comfort on leftover markup is neutral: unused structured data causes no problems, it just has no visible effect. The real case for keeping FAQPage is the extraction logic above: a clean question-answer block is the easiest format for AI systems to lift, and the Williams-Cook test shows the text inside your markup does reach the models. Add it for genuine FAQ content, never for the SERP feature. HowTo markup sits in the same bucket: its rich result died in 2023, and the structured steps stay cheap to keep where content is genuinely procedural. ### Product and Review, Only Where Real If you sell, structured price, availability, and rating data gives an AI shopping answer your numbers to quote instead of a reseller's; OpenAI's merchant feed spec and Google's Shopping surfaces both consume structured product data, which makes commerce the one corner of AI search where markup is closest to table stakes. If you do not, skip these entirely; marking up reviews that do not exist visibly on the page is the classic path to a manual action. Past that stack, niche types earn their keep only where the content genuinely exists: VideoObject with a transcript if video matters to you, SoftwareApplication if you sell software, LocalBusiness if customers walk through a door. Notice what is absent from all of this: nothing AI-specific, because per Google, no such type exists. ## How to Implement It: JSON-LD, @id, and the Entity Graph Use JSON-LD over microdata or RDFa. [Google recommends it](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data) explicitly, it lives in one script block instead of being threaded through your HTML attributes, and it is the format every generator and plugin outputs by default. Where the code physically goes, by stack: on a hand-built or framework site, a `