What AI assistants actually see of your site: the ten-minute test
A six-step test, about ten minutes, that shows you what AI crawlers really receive from your site, with every source linked and dated.
Published 2026-08-15 by SMARTCREA — 10 min read.
You asked ChatGPT to recommend a company like yours. It named three. Yours was not one of them.
The question that follows is always the same. Is my site bad, or does the assistant simply never see it? Those are two different problems with two very different price tags. The second one is often fixed in an afternoon of configuration work. The first takes months, and no amount of engineering shortens it. So the useful thing to establish first is which of the two you actually have, and you can do that yourself in about ten minutes.
Most of the advice written on this subject is editorial. Sharpen your headings. Structure your pages. Answer the questions your customers ask. All of that is true, and all of it is worthless while the crawler is receiving a blank page. This article covers only the mechanical half, the part you can measure.
The technical fact that rarely makes it into the sales deck
Your site was probably built with a modern tool. Webflow, Wix, Shopify, WordPress with a recent theme, or a React application like the one this page sits inside. In a lot of those cases the server does not send a page at all. It sends a program, and the page gets assembled inside the visitor's browser when that program runs.
Google handles this. The crawlers behind AI assistants do not.
The only serious public measurement anyone has published comes from Vercel, in a piece called "The rise of the AI crawler", dated 17 December 2024, by Giacomo Zecchini, Alice Alexandra Moore, Malte Ubl and Ryan Siddle. The conclusion leaves no room:
The results consistently show that none of the major AI crawlers currently render JavaScript.
The crawlers covered come close to the whole market: OAI-SearchBot, ChatGPT-User and GPTBot from OpenAI, ClaudeBot from Anthropic, PerplexityBot, Bytespider from ByteDance, and Meta-ExternalAgent. One distinction in that data gets lost every time the finding is repeated. They do download your JavaScript files. They never run them.
The data indicates that while ChatGPT and Claude crawlers do fetch JavaScript files (ChatGPT: 11.50%, Claude: 23.84% of requests), they don't execute them.
Google's own developer documentation, updated 4 March 2026, says something close:
Keep in mind that server-side or pre-rendering is still a great idea because it makes your website faster for users and crawlers, and not all bots can run JavaScript.
A word about how old that measurement is. Vercel ran it in December 2024. That is more than twenty months ago, and nobody has publicly reproduced it since, not Vercel and not an independent third party. The 2025 and 2026 articles that state it as current fact, including Vercel's own, all restate that single measurement. It is still the best evidence available. It is not a permanent law of the web, and that is one more reason to run the test on your own site instead of taking our word for any of this.
The six-step test
Each step stands on its own. Run them in order anyway, because the first two settle most cases.
Step 1. Read your robots.txt. One minute.
Open yourdomain.com/robots.txt in a browser. Search the text for GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot. If any of them appears above a Disallow: /, you closed the door on yourself. Usually without knowing it. A security plugin or a hosting default did it for you, and nobody sent you a note.
The trap here is assuming these settings are connected. They are not, and OpenAI states it plainly in its bots documentation:
Each setting is independent of the others – for example, a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI's generative AI foundation models.
Blocking GPTBot does not remove you from ChatGPT when search is turned on. OAI-SearchBot governs that. Plenty of companies blocked the first one believing they were protecting their content, then wondered why they were still showing up in answers. They had pulled the wrong lever, in one direction or the other.
Worth noting on the way past: being excluded from OAI-SearchBot does not erase you either. OpenAI says the affected sites "can still appear as navigational links".
Step 2. Look at the source. One minute.
This is the most useful test on the list, and it needs no tools at all.
Open your home page. Press Ctrl + U, or Cmd + Option + U on a Mac. The browser shows you the raw source, which is exactly what the server sent before anything executed. Now press Ctrl + F and search for your phone number, or for a full sentence from the page.
Find it? Your page is delivered as complete HTML, and crawlers can read it.
Find nothing but a handful of tags and an empty <div id="root"></div>? Then your page is assembled in the browser. To a crawler that does not run JavaScript, your site is blank. You can have the best content in your industry. It does not exist.
Step 3. Ask as the crawler. Three minutes.
Step 2 shows what an ordinary browser receives. Some servers, firewalls and content delivery networks answer differently depending on who is asking. To rule that out, you have to knock on the door wearing the crawler's name.
Two free services do this with nothing to install:
reqbin.com. Paste your URL, add aUser-Agent:line in the Headers section with the crawler's string, and send. No account needed.technicalseo.com/tools/fetch-render/. Easier to read if you are not technical, because it shows the received source and the final rendering side by side, so the gap between the two is visible at a glance.
The string to use for OpenAI, as published by OpenAI:
Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot
What you are looking for in the response is what you looked for in step 2: your text, or the absence of it. Plus a 200 status code. A 403 means something is actively blocking the crawler. Most often that is an anti-bot layer installed by your host, which cannot tell an AI crawler apart from a content scraper.
Step 4. Common Crawl. Two minutes.
Common Crawl is a public archive of the web, used as raw material across the industry for years. Its index is open at index.commoncrawl.org. Pick an archive from the list, type in your domain, and look at what comes back. The most recent one at the time of writing is CC-MAIN-2026-30, collected between 7 and 25 July 2026 and published on 28 July.
Now the part that gets misread, and it is misread badly in most articles on the subject. Finding nothing here proves nothing. Common Crawl says so itself, in its own FAQ:
Common Crawl's dataset is a sample of the web, and we do not generally archive any entire website but a randomly selected subset of it.
Finding your pages is a positive signal. Failing to find them is not a diagnosis. Treat this step as a bonus rather than as evidence.
Step 5. Search Console. Two minutes.
For a long time there was no way to measure your presence inside Google's generated answers. That changed on 3 June 2026, when Google opened a dedicated performance report for generative AI in Search. It covers AI Overviews and AI Mode.
Two limits to know before you go hunting for it. It is not open to everyone: Google describes the rollout as reaching "a subset of website owners". And impressions are the only metric Google documents for it. Do not plan on reading your clicks there until you have seen your clicks there with your own eyes.
So an empty report is not automatically bad news. It may only mean the report has not reached your property yet.
Step 6. Ask the question. One minute.
The last test is the simplest one. Reading it correctly is the hard part.
In ChatGPT, turn on web search and ask what a customer would ask. Which accountant, contractor, clinic or agency would you recommend in this city? Ask the same thing in Perplexity, and in Google's AI Mode.
Then ignore the answer and read the cited sources. Everything useful is in there. If the assistant cites three directories and no company sites, nobody in your market is visible and the field is wide open. If it cites your competitors, your problem stopped being technical: they have something you do not. If it cites you and then summarizes you wrong, what you have is a clarity problem rather than a visibility one.
What gets repeated, and what the sources say
| What you read everywhere | What the official source says |
|---|---|
"Block anthropic-ai in your robots.txt" | That identifier no longer appears in Anthropic's documentation. The three agents documented today are ClaudeBot, Claude-User and Claude-SearchBot. |
"Blocking Google-Extended removes you from AI Overviews" | False. Google writes: "Google-Extended does not impact a site's inclusion in Google Search nor is it used as a ranking signal in Google Search." A separate control governs AI Overviews and AI Mode. |
"Add an llms.txt file so the AI can understand you" | Google rejects this outright in its guide, updated 10 July 2026: "You don't need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search", and "Google Search ignores them". |
"Blocking GPTBot takes me out of ChatGPT" | No. The settings are independent, and OpenAI says so in its own documentation. OAI-SearchBot is what governs whether you appear in answers that use search. |
One clarification on llms.txt, because the confusion is constant. OpenAI, Anthropic and Perplexity each publish an llms.txt file for their own documentation. Publishing that file is a different act from reading anyone else's, and none of the three states anywhere that it consults them on third-party sites. We keep one on this site. It has never produced anything we can demonstrate. It also costs nothing, and we do not bill anyone for it.
What this test will not tell you
The protocol measures access. It says nothing about your reputation, and nothing about your relevance.
A site that crawlers read perfectly can still go uncited, simply because nothing else on the web says anything about you. No reviews. No press mentions. No listing that matches itself from one directory to the next. Assistants lean heavily on what other people say about you, and only partly on what you say about yourself.
It also says nothing about what happens inside the models. Nobody outside the companies that build them knows what actually decides whether a source gets picked for an answer. Treat anyone who tells you otherwise with suspicion, and that includes us.
If the test fails
If step 2 showed you an empty page, one thing matters. Your server has to send a finished page instead of a program to run. Depending on your stack, that is called server-side rendering, static generation or prerendering. It is a configuration change, and it is not a rebuild. Your design stays. Your copy stays. Your URLs stay.
Which is why the article you are reading is served as static HTML while the rest of this site is a React application. We could not write this text in good conscience and then hide it from the crawlers it describes.
Sources, all checked 15 August 2026. OpenAI bots documentation · Anthropic on web crawling · Google, common crawlers · Google, generative AI controls in Search · Google, JavaScript SEO basics · Google, AI optimization guide · Google, generative AI performance report · Vercel, The rise of the AI crawler · Common Crawl public index