{"id":29909,"date":"2026-09-28T10:34:05","date_gmt":"2026-09-28T08:34:05","guid":{"rendered":"https:\/\/wooacademy.sk\/?p=29909"},"modified":"2026-09-28T10:34:09","modified_gmt":"2026-09-28T08:34:09","slug":"technical-seo-for-ai-search","status":"publish","type":"post","link":"https:\/\/wooacademy.sk\/en\/technical-seo-for-ai-search\/","title":{"rendered":"Technical SEO for AI Search: What to Fix So ChatGPT, Claude and Google Can Read and Cite Your Site"},"content":{"rendered":"\n<p class=\"wp-block-paragraph\">A situation I see on online stores more often than you would expect: the owner asks why ChatGPT or Perplexity never mentions the brand, even though the site ranks on page one in Google. They expect the fix to be a new type of content or a secret file made for AI. Then we open the raw source of a product page and find that the price, the description and the reviews only arrive through JavaScript. Googlebot copes with that. Most AI crawlers see the header, the menu and the footer, and nothing else.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Technical SEO for AI search<\/strong> is not a new discipline. It is classic technical SEO with more visitors at the door, each with different abilities and rules. Below I cover what actually decides whether an assistant can reach your content, what the documentation from Google, OpenAI and Anthropic confirms, and which popular tips are closer to marketing than to fact. The content and brand side is covered in my <a href=\"https:\/\/wooacademy.sk\/en\/aeo-strategy-get-recommended-by-ai\/\">AEO strategy for getting recommended by AI<\/a>. Here I stay with the technical layer.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"how-ai-assistants-reach-your-website\">How AI assistants reach your website<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Your content reaches an AI answer through three routes, and something different can break on each one.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The first route is Google.<\/strong> AI Overviews and AI Mode have no separate crawler, they work with the Google Search index. The <a href=\"https:\/\/developers.google.com\/search\/docs\/appearance\/ai-features\" target=\"_blank\" rel=\"noopener\">Google documentation on AI features in Search<\/a> says a page must be indexed and eligible to show with a snippet, with no additional requirements. Healthy indexing is the entry ticket to AI Overviews.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The second route is the search indexes of AI companies.<\/strong> OpenAI runs OAI-SearchBot, Anthropic runs Claude-SearchBot and Perplexity runs PerplexityBot. When an assistant answers with web search, it picks sources from these indexes.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>The third route is real-time fetching.<\/strong> When a user pastes a link or the assistant needs one specific page, it sends an agent such as ChatGPT-User or Claude-User. The agent downloads the page on the spot and waits neither for a slow server nor for scripts.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Training crawlers such as GPTBot and ClaudeBot run alongside, collecting data for future models. They have no direct effect on whether an assistant cites you in search today. Mixing up these categories is the most common mistake I find in audits.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"which-ai-crawlers-to-allow-and-which-you-can-block\">Which AI crawlers to allow and which you can block<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The <a href=\"https:\/\/en.wikipedia.org\/wiki\/Robots.txt\" target=\"_blank\" rel=\"noopener\">robots.txt<\/a> file is still the main way to tell crawlers where they may go. Yet many site owners paste ready-made blocks into it without knowing what each bot does.<\/p>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Crawler<\/th><th>Operator<\/th><th>Purpose<\/th><th>What blocking it does<\/th><\/tr><\/thead><tbody><tr><td>Googlebot<\/td><td>Google<\/td><td>index for Search, AI Overviews and AI Mode<\/td><td>you drop out of Search and its AI features<\/td><\/tr><tr><td>Google-Extended<\/td><td>Google, robots.txt token only<\/td><td>Gemini training and grounding in Gemini Apps and Vertex AI<\/td><td>no effect on Google Search or rankings<\/td><\/tr><tr><td>OAI-SearchBot<\/td><td>OpenAI<\/td><td>ChatGPT search<\/td><td>you do not appear in ChatGPT search results<\/td><\/tr><tr><td>GPTBot<\/td><td>OpenAI<\/td><td>model training<\/td><td>content is not used for training, search stays<\/td><\/tr><tr><td>ChatGPT-User<\/td><td>OpenAI<\/td><td>opening a page at user request<\/td><td>robots.txt may not apply to it<\/td><\/tr><tr><td>Claude-SearchBot<\/td><td>Anthropic<\/td><td>index for Claude search<\/td><td>lower visibility in Claude results<\/td><\/tr><tr><td>ClaudeBot<\/td><td>Anthropic<\/td><td>model training<\/td><td>content is excluded from future training data<\/td><\/tr><tr><td>Claude-User<\/td><td>Anthropic<\/td><td>opening a page at user request<\/td><td>Claude cannot open the page when asked about it<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">The <a href=\"https:\/\/developers.openai.com\/api\/docs\/bots\" target=\"_blank\" rel=\"noopener\">OpenAI documentation on its crawlers<\/a> gives a simple rule: to appear in ChatGPT search results, allow OAI-SearchBot. GPTBot is a separate decision about training, so you can opt out of training and stay in search. Anthropic uses the <a href=\"https:\/\/support.claude.com\/en\/articles\/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler\" target=\"_blank\" rel=\"noopener\">same three-bot split<\/a> and also respects Crawl-delay.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The most frequent misunderstanding is Google-Extended. It is not a crawler but a robots.txt token that controls use of your content for Gemini. The <a href=\"https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/google-common-crawlers\" target=\"_blank\" rel=\"noopener\">Google overview of its common crawlers<\/a> states that it does not affect inclusion in Google Search and is not a ranking signal. Blocking it will not remove you from AI Overviews.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For a typical company site or online store, I recommend allowing the search and user-triggered bots and deciding about training bots by business model. A publisher living off its content has different interests from a store that wants models to know its brand. EU publishers have a legal angle too: the commercial text and data mining exception in the <a href=\"https:\/\/en.wikipedia.org\/wiki\/Directive_on_Copyright_in_the_Digital_Single_Market\" target=\"_blank\" rel=\"noopener\">Directive on Copyright in the Digital Single Market<\/a> applies only if the rightsholder has not opted out.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"cloudflare-hosting-and-firewalls-as-silent-blockers\">Cloudflare, hosting and firewalls as silent blockers<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Robots.txt is only half the story. Cloudflare offers a one-click Block AI Bots setting and a managed robots.txt that adds disallow rules for GPTBot, ClaudeBot, Google-Extended and other bots. According to the <a href=\"https:\/\/blog.cloudflare.com\/content-independence-day-ai-options\/\" target=\"_blank\" rel=\"noopener\">Cloudflare announcement<\/a>, from September 15, 2026 new domains default to allowing search bots while blocking training and agent bots on pages that display ads. An ad-supported blog on a new domain can therefore stop an assistant from opening the very article a user asks about.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">WordPress security plugins and hosting firewalls add more layers. Each can return a 403 or 429 error to an AI crawler without any trace in robots.txt. The only reliable place to see what really happens is the server access log.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"javascript-is-the-biggest-blind-spot-of-ai-crawlers\">JavaScript is the biggest blind spot of AI crawlers<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In my experience this is the most common reason a site is invisible to AI. A <a href=\"https:\/\/vercel.com\/blog\/the-rise-of-the-ai-crawler\" target=\"_blank\" rel=\"noopener\">Vercel analysis of AI crawler traffic<\/a>, covering roughly 1.3 billion monthly AI bot requests across its network, found that the crawlers from OpenAI, Anthropic and Perplexity do not execute JavaScript. They download JS files but never run them. The exceptions are Gemini, which uses Googlebot infrastructure, and AppleBot.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Googlebot renders JavaScript in headless Chromium, though it queues pages for rendering first. Even so, the Google guide to <a href=\"https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/javascript\/javascript-seo-basics\" target=\"_blank\" rel=\"noopener\">JavaScript SEO basics<\/a> recommends server-side rendering or prerendering, because not every bot can run JavaScript.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The same study holds another useful number. The ChatGPT and Claude crawlers sent about 34% of requests to pages that no longer exist, versus roughly 8% for Googlebot. AI bots often try old URLs, so clean redirects after a migration matter even more for AI than for Google.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">On online stores, these are the elements I most often find missing from the HTML:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>price and availability on variable products<\/li>\n\n\n<li>reviews injected by a third-party widget script<\/li>\n\n\n<li>description or specification tabs loaded only after a click<\/li>\n\n\n<li>entire storefronts built as single-page applications without server-side rendering<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The test is simple. Open the page source, not the element inspector, and search for the price or the first sentence of the description. If it is not there, an AI crawler does not see it.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"indexing-and-snippets-are-your-ticket-into-ai-overviews\">Indexing and snippets are your ticket into AI Overviews<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Google pulls content for its AI features from the regular index, so anything that hurts indexing hurts AI visibility too: canonical tags pointing to the wrong URL, duplicates created by filters, thousands of useless URLs from internal site search. If Google takes weeks to index new pages, see my comparison of the <a href=\"https:\/\/wooacademy.sk\/en\/best-google-indexing-tools\/\">best Google indexing tools<\/a>, and before paying for a service that promises results, read <a href=\"https:\/\/wooacademy.sk\/en\/how-guaranteed-indexing-tools-work\/\">how guaranteed indexing tools work<\/a>.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A less known trap is snippet restrictions. A nosnippet directive, or max-snippet set to zero, removes you from AI Overviews as well, as described in the <a href=\"https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/robots-meta-tag\" target=\"_blank\" rel=\"noopener\">robots meta tag documentation<\/a>. To keep only part of a page out of answers, use the data-nosnippet attribute on that element.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"speed-and-core-web-vitals-facts-versus-guesses\">Speed and Core Web Vitals: facts versus guesses<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">The thresholds are clear. According to <a href=\"https:\/\/web.dev\/articles\/vitals\" target=\"_blank\" rel=\"noopener\">web.dev<\/a>, a good LCP is 2.5 seconds or less, INP 200 milliseconds or less and CLS 0.1 or less, measured at the 75th percentile of real visits.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The claim that AI crawlers skip slow sites and take information from faster competitors is not documented anywhere, though. Neither OpenAI nor Anthropic publishes crawler timeouts, and LCP, INP and CLS measure the experience of a human, not a bot.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What is documented is server response. The Google guide to <a href=\"https:\/\/developers.google.com\/search\/docs\/crawling-indexing\/large-site-managing-crawl-budget\" target=\"_blank\" rel=\"noopener\">managing crawl budget<\/a> explains that crawl capacity rises when a site responds quickly and consistently, and drops with slowdowns or 5xx and 429 errors. So time to first byte and server stability matter more for AI than a PageSpeed score. For real-time agents, I treat a response of several seconds as a risk, even without a hard number.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Speed is worth fixing anyway for people and conversions, usually starting with third-party scripts.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"page-structure-that-both-machines-and-people-can-read\">Page structure that both machines and people can read<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Once a crawler has the HTML, it still needs to find the main content. Ordinary things help:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>one H1 and a logical H2 and H3 hierarchy without skipped levels<\/li>\n\n\n<li>semantic main, article and nav elements that separate content from menus and footers<\/li>\n\n\n<li>tables as real HTML tables, not images or PDFs<\/li>\n\n\n<li>key facts in the text, not only in a banner or video<\/li>\n\n\n<li>breadcrumbs and internal links between related topics<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">The Google <a href=\"https:\/\/developers.google.com\/search\/docs\/fundamentals\/ai-optimization-guide\" target=\"_blank\" rel=\"noopener\">guide to optimizing for generative AI in Search<\/a> notes there is no need to break content into tiny pieces. Clean structure is not an AI trick, it is basic readability. The writing itself is covered in my guide on <a href=\"https:\/\/wooacademy.sk\/en\/how-to-write-content-ai-cites\/\">how to write content AI cites<\/a>.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"structured-data-useful-but-not-magic\">Structured data: useful, but not magic<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In the same guide, Google states that structured data is not required for generative AI search and no special schema exists for it. The same goes for llms.txt and other AI files, which Google Search does not use.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">I still implement schema, just for different reasons than the ones usually sold. Product, Offer, Organization, Article and BreadcrumbList types from <a href=\"https:\/\/schema.org\/\" target=\"_blank\" rel=\"noopener\">schema.org<\/a> still unlock rich results in classic search, and they force you to keep facts consistent. If a product shows one price in the schema and another in the HTML or the merchant feed, you have a data trust problem regardless of AI.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"how-to-measure-ai-visibility\">How to measure AI visibility<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In 2026 Search Console added a <a href=\"https:\/\/support.google.com\/webmasters\/answer\/16984139\" target=\"_blank\" rel=\"noopener\">performance report for generative AI features<\/a>. It shows impressions in AI Overviews and AI Mode by page, country and device. It does not include clicks yet, and a property needs enough impressions before data appears.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Google also added a setting that excludes a whole site from generative AI features in Search. Regular results stay, but you lose impressions in AI Overviews, AI Mode and AI elements in Discover. It rarely makes sense for stores, but it is far more precise than blocking Google-Extended.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">For other assistants you have server logs and analytics, where ChatGPT traffic carries the utm_source parameter set to chatgpt.com.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"popular-advice-that-does-not-hold-up\">Popular advice that does not hold up<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Some widely repeated claims I would not put my name under.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A 30% performance lift thanks to schema.<\/strong> The figure comes from one vendor, with no published methodology, no control group and no definition of AI performance. Google says schema is not required for its AI features.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>AI crawlers skip slow sites.<\/strong> Logical, but undocumented. Far more often I have seen a bot fail to get past a firewall or JavaScript.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>An audit in a popular SEO tool is enough.<\/strong> SEO tool crawlers usually behave like Googlebot or a browser with JavaScript on. Until you switch rendering off, they show a site GPTBot will never see.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A monthly audit for everyone.<\/strong> Problems appear with a migration, a theme change, a plugin update or new hosting protection, not on a calendar. I have watched a brand lose about four fifths of its organic traffic after a redesign because nobody checked the technical basics afterwards.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>A case study without data.<\/strong> Growing from 10% to 30% of keywords with AI Overviews through technical fixes alone is possible, especially after broken indexing. Without a baseline, it is a story, not evidence.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What these guides rarely say is that technical SEO is only the gate. It decides whether AI can reach you. Whether AI cites you depends on your content and on what others say about you, which is why <a href=\"https:\/\/wooacademy.sk\/en\/brand-mentions-ai-citations\/\">brand mentions and AI citations<\/a> deserve as much attention as your robots.txt.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"technical-seo-audit-for-ai-search-step-by-step\">Technical SEO audit for AI search step by step<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">This is how I run an audit focused on AI visibility. With access to hosting and Search Console, you can do it yourself.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"1-check-robots-txt-and-cdn-settings\">1. Check robots.txt and CDN settings<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Go through every User-agent block and confirm that Googlebot, OAI-SearchBot, Claude-SearchBot and PerplexityBot can reach important sections. Then review the AI bot settings in Cloudflare or your CDN.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"2-verify-in-the-logs-what-ai-crawlers-receive\">2. Verify in the logs what AI crawlers receive<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Pull at least one week of access logs and filter for AI bots. Repeated 403 or 429 responses mean blocking, a high share of 404s points to missing redirects.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"3-compare-the-source-html-with-what-a-visitor-sees\">3. Compare the source HTML with what a visitor sees<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">For the homepage, a category, a product and an article, look for key information in the page source or load the page with JavaScript disabled. Whatever disappears is invisible to AI crawlers.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"4-review-indexing-and-snippets-in-search-console\">4. Review indexing and snippets in Search Console<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">In the Pages report, find important URLs that are not indexed and check the canonicals Google selected. Use URL Inspection to confirm there is no nosnippet or zero max-snippet directive.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"5-measure-server-response-and-core-web-vitals\">5. Measure server response and Core Web Vitals<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Check the Core Web Vitals report and the average response time in Crawl stats. If response times fluctuate or 5xx errors grow, fix hosting and caching first.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"6-check-html-structure-and-internal-links\">6. Check HTML structure and internal links<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Run a crawler with JavaScript rendering off and verify headings, click depth and orphaned URLs.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"7-validate-structured-data-and-fact-consistency\">7. Validate structured data and fact consistency<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Test the main templates in the Rich Results Test and compare price, availability and name in the schema with the visible text and the merchant feed.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"8-set-up-ongoing-monitoring\">8. Set up ongoing monitoring<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Turn on Search Console alerts, follow the generative AI features report and repeat steps 1 to 3 after every migration or security change. If your team ships code with AI assistants, the checks in my guide to <a href=\"https:\/\/wooacademy.sk\/en\/?p=29915\">secure AI coding for WordPress and WooCommerce<\/a> help keep a quick fix from breaking rendering or crawler access.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"technical-seo-for-ai-search-checklist\">Technical SEO for AI search checklist<\/h2>\n\n\n\n<figure class=\"wp-block-table\"><table><thead><tr><th>Area<\/th><th>What to verify<\/th><th>How<\/th><th>Priority<\/th><\/tr><\/thead><tbody><tr><td>Crawler access<\/td><td>AI search and user bots are not blocked<\/td><td>robots.txt, CDN, firewall<\/td><td>high<\/td><\/tr><tr><td>Server responses<\/td><td>AI bots do not get 403, 429 or 5xx<\/td><td>access log<\/td><td>high<\/td><\/tr><tr><td>JavaScript<\/td><td>price, description and main text are in the source HTML<\/td><td>page source, JavaScript off<\/td><td>high<\/td><\/tr><tr><td>Indexing and snippets<\/td><td>pages indexed without nosnippet<\/td><td>Search Console, URL Inspection<\/td><td>high<\/td><\/tr><tr><td>Redirects<\/td><td>old URLs return 301 to the right pages<\/td><td>404 log, crawler<\/td><td>medium<\/td><\/tr><tr><td>Server response time<\/td><td>stable response time<\/td><td>Crawl stats<\/td><td>medium<\/td><\/tr><tr><td>Core Web Vitals<\/td><td>LCP up to 2.5 s, INP up to 200 ms, CLS up to 0.1<\/td><td>Core Web Vitals report<\/td><td>medium<\/td><\/tr><tr><td>HTML structure<\/td><td>one H1, logical headings, HTML tables<\/td><td>crawler, template review<\/td><td>medium<\/td><\/tr><tr><td>Structured data<\/td><td>valid schema consistent with the text<\/td><td>Rich Results Test<\/td><td>medium<\/td><\/tr><tr><td>Measurement<\/td><td>impressions in AI features<\/td><td>generative AI features report<\/td><td>low<\/td><\/tr><\/tbody><\/table><\/figure>\n\n\n\n<p class=\"wp-block-paragraph\">If you would rather hand this list to someone, WooAcademy (Milan Fra\u0148o) works remotely with companies in the EU and the US on technical SEO audits and AI search visibility audits, including WooCommerce stores, where rendering and crawler access tend to be the weak spots.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\" id=\"faq\">FAQ<\/h2>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"do-i-need-an-llms-txt-file-for-ai-to-cite-my-site\">Do I need an llms.txt file for AI to cite my site?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">No. Google states in its guide to optimizing for generative AI in Search that no new machine-readable or AI text files are needed to appear in its AI features, and Google Search does not use llms.txt. Your time is better spent getting the main content into the HTML and letting search and user-triggered crawlers in.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"if-i-block-gptbot-will-i-disappear-from-chatgpt\">If I block GPTBot, will I disappear from ChatGPT?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Not from ChatGPT search. According to OpenAI, GPTBot collects data for model training, while OAI-SearchBot powers ChatGPT search results. Keep OAI-SearchBot allowed and you stay visible, while blocking GPTBot only excludes your content from training future models.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"does-blocking-google-extended-affect-my-rankings-or-ai-overviews\">Does blocking Google-Extended affect my rankings or AI Overviews?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">No. Google-Extended controls use of your content for Gemini Apps and Vertex AI, and Google says it does not affect Search inclusion or rankings. AI Overviews and AI Mode use the regular Search index. To opt out of them, use the generative AI features setting in Search Console or snippet restrictions.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"can-ai-crawlers-see-content-loaded-with-javascript\">Can AI crawlers see content loaded with JavaScript?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Mostly not. The Vercel analysis found that crawlers from OpenAI, Anthropic and Perplexity do not execute JavaScript, while Googlebot and Gemini do render it. Price, availability, descriptions and main article text should therefore be in the HTML your server sends.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"is-schema-markup-required-to-appear-in-ai-overviews\">Is schema markup required to appear in AI Overviews?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">No. Google states that structured data is not required for generative AI in Search and no special schema exists. Schema still helps with rich results and with keeping product and company data consistent, so it is worth having, just not as a magic lever for AI.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"how-can-i-tell-whether-my-site-appears-in-ai-overviews-and-ai-mode\">How can I tell whether my site appears in AI Overviews and AI Mode?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Google Search Console has a performance report for generative AI features showing impressions in AI Overviews and AI Mode by page, country and device, though not clicks yet. For ChatGPT, Claude and Perplexity, use server logs and referral traffic in analytics, where ChatGPT visits carry the chatgpt.com utm_source value.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\" id=\"how-often-should-i-run-a-technical-audit-focused-on-ai\">How often should I run a technical audit focused on AI?<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">A thorough quarterly review is enough for a smaller site, while a larger store with frequent changes benefits from monthly checks. More important is repeating the robots.txt, firewall and rendering checks after every migration, theme change or hosting change, because that is when most problems appear.<\/p>\n\n\n\n<script type=\"application\/ld+json\">{\"@context\": \"https:\/\/schema.org\", \"@graph\": [{\"@type\": \"FAQPage\", \"@id\": \"https:\/\/wooacademy.sk\/en\/technical-seo-for-ai-search\/#faq\", \"mainEntity\": [{\"@type\": \"Question\", \"name\": \"Do I need an llms.txt file for AI to cite my site?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"No. Google states in its guide to optimizing for generative AI in Search that no new machine-readable or AI text files are needed to appear in its AI features, and Google Search does not use llms.txt. Your time is better spent getting the main content into the HTML and letting search and user-triggered crawlers in.\"}}, {\"@type\": \"Question\", \"name\": \"If I block GPTBot, will I disappear from ChatGPT?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"Not from ChatGPT search. According to OpenAI, GPTBot collects data for model training, while OAI-SearchBot powers ChatGPT search results. Keep OAI-SearchBot allowed and you stay visible, while blocking GPTBot only excludes your content from training future models.\"}}, {\"@type\": \"Question\", \"name\": \"Does blocking Google-Extended affect my rankings or AI Overviews?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"No. Google-Extended controls use of your content for Gemini Apps and Vertex AI, and Google says it does not affect Search inclusion or rankings. AI Overviews and AI Mode use the regular Search index. To opt out of them, use the generative AI features setting in Search Console or snippet restrictions.\"}}, {\"@type\": \"Question\", \"name\": \"Can AI crawlers see content loaded with JavaScript?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"Mostly not. The Vercel analysis found that crawlers from OpenAI, Anthropic and Perplexity do not execute JavaScript, while Googlebot and Gemini do render it. Price, availability, descriptions and main article text should therefore be in the HTML your server sends.\"}}, {\"@type\": \"Question\", \"name\": \"Is schema markup required to appear in AI Overviews?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"No. Google states that structured data is not required for generative AI in Search and no special schema exists. Schema still helps with rich results and with keeping product and company data consistent, so it is worth having, just not as a magic lever for AI.\"}}, {\"@type\": \"Question\", \"name\": \"How can I tell whether my site appears in AI Overviews and AI Mode?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"Google Search Console has a performance report for generative AI features showing impressions in AI Overviews and AI Mode by page, country and device, though not clicks yet. For ChatGPT, Claude and Perplexity, use server logs and referral traffic in analytics, where ChatGPT visits carry the chatgpt.com utm_source value.\"}}, {\"@type\": \"Question\", \"name\": \"How often should I run a technical audit focused on AI?\", \"acceptedAnswer\": {\"@type\": \"Answer\", \"text\": \"A thorough quarterly review is enough for a smaller site, while a larger store with frequent changes benefits from monthly checks. More important is repeating the robots.txt, firewall and rendering checks after every migration, theme change or hosting change, because that is when most problems appear.\"}}]}, {\"@type\": \"HowTo\", \"name\": \"How to run a technical SEO audit for AI search\", \"step\": [{\"@type\": \"HowToStep\", \"position\": 1, \"name\": \"Check robots.txt and CDN settings\", \"text\": \"Go through every User-agent block and confirm that Googlebot, OAI-SearchBot, Claude-SearchBot and PerplexityBot can reach important sections. Then review the AI bot settings in Cloudflare or your CDN.\"}, {\"@type\": \"HowToStep\", \"position\": 2, \"name\": \"Verify in the logs what AI crawlers receive\", \"text\": \"Pull at least one week of access logs and filter for AI bots. Repeated 403 or 429 responses mean blocking, a high share of 404s points to missing redirects.\"}, {\"@type\": \"HowToStep\", \"position\": 3, \"name\": \"Compare the source HTML with what a visitor sees\", \"text\": \"For the homepage, a category, a product and an article, look for key information in the page source or load the page with JavaScript disabled. Whatever disappears is invisible to AI crawlers.\"}, {\"@type\": \"HowToStep\", \"position\": 4, \"name\": \"Review indexing and snippets in Search Console\", \"text\": \"In the Pages report, find important URLs that are not indexed and check the canonicals Google selected. Use URL Inspection to confirm there is no nosnippet or zero max-snippet directive.\"}, {\"@type\": \"HowToStep\", \"position\": 5, \"name\": \"Measure server response and Core Web Vitals\", \"text\": \"Check the Core Web Vitals report and the average response time in Crawl stats. If response times fluctuate or 5xx errors grow, fix hosting and caching first.\"}, {\"@type\": \"HowToStep\", \"position\": 6, \"name\": \"Check HTML structure and internal links\", \"text\": \"Run a crawler with JavaScript rendering off and verify headings, click depth and orphaned URLs.\"}, {\"@type\": \"HowToStep\", \"position\": 7, \"name\": \"Validate structured data and fact consistency\", \"text\": \"Test the main templates in the Rich Results Test and compare price, availability and name in the schema with the visible text and the merchant feed.\"}, {\"@type\": \"HowToStep\", \"position\": 8, \"name\": \"Set up ongoing monitoring\", \"text\": \"Turn on Search Console alerts, follow the generative AI features report and repeat steps 1 to 3 after every migration or security change. If your team ships code with AI assistants, the checks in my guide to secure AI coding for WordPress and WooCommerce help keep a quick fix from breaking rendering or crawler access.\"}], \"@id\": \"https:\/\/wooacademy.sk\/en\/technical-seo-for-ai-search\/#howto\", \"inLanguage\": \"en\"}, {\"@type\": \"Person\", \"@id\": \"https:\/\/wooacademy.sk\/#milan-frano\", \"name\": \"Milan Fra\u0148o\", \"url\": \"https:\/\/wooacademy.sk\/\", \"worksFor\": {\"@type\": \"Organization\", \"name\": \"wooacademy.sk\", \"url\": \"https:\/\/wooacademy.sk\/\"}, \"knowsAbout\": [\"SEO\", \"Technick\u00e9 SEO\", \"WordPress\", \"Google Search Console\"]}]}<\/script>\n\n\n<!-- wp:themify-builder\/canvas \/-->","protected":false},"excerpt":{"rendered":"<p>The technical factors that decide whether AI assistants can read your site at all. Crawler access, JavaScript, indexing, speed and schema without the myths, backed by Google, OpenAI and Anthropic documentation.<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[824],"tags":[828],"class_list":["post-29909","post","type-post","status-publish","format-standard","hentry","category-seo","tag-seo-doktor","has-post-title","has-post-date","has-post-category","has-post-tag","has-post-comment","has-post-author",""],"builder_content":"","_links":{"self":[{"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/posts\/29909","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/comments?post=29909"}],"version-history":[{"count":0,"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/posts\/29909\/revisions"}],"wp:attachment":[{"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/media?parent=29909"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/categories?post=29909"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/wooacademy.sk\/en\/wp-json\/wp\/v2\/tags?post=29909"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}