In plain terms
Search engines have sent robots to read web pages for decades. AI companies now send their own, for three different jobs: gathering text to train future models, indexing pages so they can be shown as sources in AI search, and fetching one page at the moment a user asks about it. Each robot announces itself with a name, such as GPTBot or ClaudeBot.
Why it matters
You decide which of these may read your site, and the decision has consequences in both directions. Block everything and you disappear from AI answers. Allow everything and your content may be used for training without any return. Because the three jobs use different robots, you can choose: many publishers block training and allow search. Check your current settings, since some hosting and security services block AI robots by default.
Example
A software company ranks well on Google and is never cited by ChatGPT. The cause is a robots.txt file written in 2023, during a wave of concern about training, which blocks every AI robot including the search one. The company keeps the training robot blocked and allows the search and user-request robots. A few weeks later its pages begin to appear as sources.
Most often confused with
AI Crawler vs. Search engine crawler
The technology is the same; the bargain is different. A search crawler took your content and sent visitors back. An AI crawler may take your content and return an answer that needs no visit. That is why site owners look at AI crawlers more critically, and why access is now set per purpose.
Under the hood
Operators publish separate user agents per purpose. OpenAI: GPTBot (training), OAI-SearchBot (search index), ChatGPT-User (user-initiated fetches). Anthropic: ClaudeBot, Claude-SearchBot, Claude-User. Perplexity: PerplexityBot and Perplexity-User. Google reads pages with Googlebot for Search and its AI features alike; Google-Extended is a robots.txt token that controls use for Gemini, and no separate crawler. Others include Applebot-Extended, CCBot (Common Crawl) and Meta-ExternalAgent. Controls: robots.txt, which is a voluntary standard; firewall or CDN rules; and verification against published IP ranges, because a user-agent string can be faked. Policies for user-initiated fetchers differ between operators. Many AI crawlers do not run JavaScript, so content that appears only after scripts load may be invisible to them. Names and rules change: check each operator's documentation.