Answer PatchAnswer Patch

Why ChatGPT and Perplexity Can't Find You

AI assistants skip businesses they can't parse. Why ChatGPT and Perplexity miss yours, and the fixes that make you legible to both.

Answer Patch Team10 min read

ChatGPT and Perplexity skip your business because they cannot parse it — not because they have not heard of it. Both systems retrieve web pages in real time when a user asks a question, read what is actually on the page, and decide in seconds whether the content answers clearly enough to cite. If your site buries what you do behind JavaScript a crawler cannot render, never states your offering in a plain sentence, or blocks the wrong crawler in robots.txt, the AI moves on to a competitor that does those things. The good news: every reason AI systems skip a business is checkable and fixable. Gartner predicted in February 2024 that traditional search engine volume would drop 25% by 2026 due to AI chatbots. Whether or not that exact number lands, the direction is clear — a growing share of buyer research now happens inside AI assistants, and the businesses those assistants can parse are the ones that get cited.

How ChatGPT and Perplexity Actually Find Websites

ChatGPT does not have its own web index. Every URL it considers for citation was first discovered through Bing. When a user triggers a search inside ChatGPT, the system queries Bing's index for candidate pages, then sends a live fetcher — identified as ChatGPT-User — to read the top results in real time. If Bing has not indexed your page, ChatGPT cannot find it. Perplexity works differently. It runs a retrieval-first architecture: every query triggers a live web search, and every answer comes with inline source citations. Perplexity maintains its own crawl index via PerplexityBot while also pulling from external search APIs. The common thread is that both systems fetch and read your actual page content at query time. They are not pulling from a static training snapshot. They are reading your site right now, deciding whether to cite it based on what they find.

Why Do AI Systems Skip Certain Businesses?

AI systems skip businesses for concrete, technical reasons — not because the business is too small or too niche. The five primary disqualifiers are: the page never states what the business does in plain language (a problem of business clarity), key content loads only via client-side JavaScript that a non-browser fetch returns as an empty div, robots.txt blocks the AI search crawler — often by accident, when the site owner meant to block only training — there is no structured data a machine can use to verify basic facts, and the content reads as marketing copy rather than direct answers to real questions. A study from Rutgers Business School and Wharton, published December 2025, found that publishers blocking AI crawlers via robots.txt experienced a 23.1% decline in total monthly visits. The penalty for being unparseable is not theoretical — it is measurable and compounds as AI search traffic grows.

SignalCited SitesSkipped Sites
Business descriptionPlain sentence stating what you do in the first 200 wordsImplied through images, taglines, or navigation labels
Content renderingServer-rendered HTML readable without JavaScriptKey content hidden behind client-side JS frameworks
robots.txt policyOAI-SearchBot and PerplexityBot explicitly allowedGPTBot blocked, accidentally blocking search crawlers too
Structured dataJSON-LD with Organization, FAQ, or Article schemaNo machine-readable schema on any page
Content formatDirect answers with H2/H3 hierarchy and listsDense marketing prose with no question-answer structure
Trust signalsReviews, credentials, or certifications visible in HTMLSocial proof buried in images or absent entirely

Parseability comes down to whether an AI system can extract a clear, verified answer from your page without guessing. When ChatGPT-User or Perplexity-User fetches a page, it reads the raw HTML — not the rendered browser view. If your primary content sits behind a JavaScript framework that requires execution to display, the crawler sees a shell. After the content is readable, the system looks for structure: heading hierarchy (H2s and H3s that break the page into scannable sections), lists, FAQ blocks, and paragraphs that each make a self-contained point. Pages with this structure get cited more than dense, unbroken prose, because the AI can extract a specific answer without risk of misrepresenting the source. That structural clarity is what answer readiness means in practice. The final layer is verification: JSON-LD schema gives the model machine-readable facts it can cross-check against the page content.

The Crawlers That Control Your AI Visibility

OpenAI and Perplexity each run multiple crawlers, and blocking the wrong one is one of the most common mistakes site owners make. OpenAI operates three: GPTBot collects content for model training, OAI-SearchBot indexes pages for ChatGPT's search feature, and ChatGPT-User fetches pages live during conversations. Blocking GPTBot stops training use but does not remove you from ChatGPT search. Blocking OAI-SearchBot does. As of December 2025, ChatGPT-User ignores robots.txt entirely — OpenAI treats it as a "technical extension of the user." Perplexity runs PerplexityBot for indexing and Perplexity-User for live fetches. A Q3 2026 crawl of 1,744 websites found that 84.2% have no AI crawler policy at all in their robots.txt — meaning most sites have not made a deliberate choice about which AI systems can access their content. That is a technical access gap with direct visibility consequences.

CrawlerOperatorPurposeWhat Blocking Does
GPTBotOpenAIModel training data collectionStops training use; does NOT affect ChatGPT search
OAI-SearchBotOpenAIChatGPT search indexRemoves your pages from ChatGPT search results
ChatGPT-UserOpenAILive page fetch during chatsIgnores robots.txt as of December 2025
PerplexityBotPerplexitySearch index maintenanceRemoves your pages from Perplexity's search index
Perplexity-UserPerplexityLive page fetch during queriesBlocks real-time citation for relevant questions

Does Structured Data Actually Help AI Citations?

The evidence is real but mixed. JSON-LD now appears on 41% of all pages, according to the HTTP Archive Web Almanac, up from 34% two years prior. An analysis of 73 websites found that those with properly implemented structured data were cited in AI responses 3.2 times more often than those without. But Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026, matched them against 4,000 control pages, and found no statistically significant uplift in citations across Google AI Overviews, AI Mode, or ChatGPT. The reconciliation is that structured data is necessary but not sufficient. It reduces ambiguity when an AI system has to choose between competing sources — SearchVIU confirmed in October 2025 that ChatGPT, Claude, Perplexity, and Gemini all actively process schema markup when accessing content. But schema on a page with weak content does not create citations from nothing.

What Is llms.txt and Should You Add One?

The llms.txt specification, proposed by Jeremy Howard in September 2024, is a Markdown file placed at your site's root that tells AI systems what your site is, which pages matter most, and how you want your content used. It functions as a curated table of contents — ideally 20 to 50 high-signal links with descriptions — that AI crawlers can read in one pass. The practical reality as of mid-2026: no major consumer AI search engine, including ChatGPT, Perplexity, Google AI Overviews, and Gemini, has publicly confirmed that it consumes llms.txt to answer user queries. Google's John Mueller explicitly stated in September 2025 that Google does not use llms.txt for search or AI Overviews. The tools that do use it today are developer-facing: Cursor, Windsurf, and documentation platforms like Mintlify. Adding an llms.txt file is low-effort and low-risk, but it is not currently a factor in whether ChatGPT or Perplexity cites your business.

Why AI Search Visitors Convert at Higher Rates

AI search traffic converts at dramatically higher rates than traditional organic search, which makes the cost of being invisible to these systems a real revenue problem. Ahrefs published their own internal data in June 2025: AI search accounted for just 0.5% of their total traffic but drove 12.1% of all signups — a 23x conversion rate compared to organic. Seer Interactive found ChatGPT referral traffic converting at 15.9% versus Google organic at 1.76% for a B2B client, with Perplexity at 10.5%. Semrush's broader cross-industry research established a 4.4x conversion baseline for AI-referred visitors. The reason is structural: AI visitors arrive pre-qualified. The AI platform has already handled the research, comparison, and narrowing stages. Users who click through from an AI citation are investigating a specific recommendation, not scanning a list of ten blue links. That pre-qualification makes every AI-referred visit disproportionately valuable.

SourceAI Search ConversionGoogle OrganicMultiple
Ahrefs internal data (June 2025)0.5% of traffic drove 12.1% of signupsBaseline organic23x
Seer Interactive, B2B client (Oct 2024–Apr 2025)ChatGPT: 15.9%, Perplexity: 10.5%1.76%~9x
Semrush cross-industry (2025)4.4x baseline conversion rateBaseline organic4.4x

Six Fixes That Make Your Business Findable

The fixes map to specific, testable changes — not vague advice about "improving your online presence." Each one addresses a reason AI systems pass over a site, and each can be checked in under ten minutes. The first three deal with what the AI can read: making sure your homepage states what you do in a plain sentence within the first 200 words, ensuring that content renders as server-side HTML rather than requiring JavaScript execution, and configuring robots.txt to allow OAI-SearchBot, ChatGPT-User, PerplexityBot, and Perplexity-User while optionally blocking GPTBot if you want to prevent training use. The second three deal with what the AI can verify: adding JSON-LD schema (Organization or LocalBusiness at minimum, FAQPage where relevant), making reviews and credentials visible in the HTML rather than in images, and structuring content with H2/H3 headings and direct answers to real buyer questions. These six areas map directly to the pillars Answer Patch checks in its free scan.

Quick Checklist: Is Your Site Parseable?

Every reason ChatGPT and Perplexity skip a business traces back to one of six measurable areas: business clarity, answer readiness, technical access, structured data, trust signals, and competitive visibility. The fixes are not speculative — they are specific changes to specific files and page structures, each one testable against real crawled evidence. A site that was invisible to AI search yesterday can be citable tomorrow, because these systems re-fetch pages in real time. Answer Patch's free scan checks all six areas in about two minutes, scoring each one and showing exactly where a site falls short. No account required, no credit card, no sales call. If the scan surfaces issues you want fixed for you, the paid Fix Report provides up to ten prioritized, copy-ready changes — but the scan itself tells you whether any of this applies to your site before you spend anything.

Why ChatGPT and Perplexity Can't Find You — Answer Patch