Key ecosystem metrics across 5,000+ websites using Agent Analytics and AI Chat Referral Tracking.
↓ 2%
Compared to the previous 90 days
The amount of visits from known agents vs. humans
↑ 7%
Compared to the previous 90 days
The percentage of bot traffic that's AI-related
↑ 15%
Compared to the previous 90 days
AI Agent
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
AI Assistant
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding AgentFetches documentation and other resources to help build software
AI Data Provider
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
AI Search Crawler
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
Archiver
ArchiverCaptures and stores historical website snapshots for long-term digital preservation
Automated Agent
Automated AgentAutomates browser interactions programmatically without direct human supervision
Developer Helper
Developer HelperAssists with testing, debugging, and ensuring website functionality
Fetcher
FetcherRetrieves web page metadata to power app features like link previews or feeds
Intelligence Gatherer
Intelligence GathererAnalyzes web content for brand safety, competitive insights, and ad targeting
Scraper
ScraperExtracts large amounts of web data, often without explicit website permission
Search Engine Crawler
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
Security Scanner
Security ScannerScans websites for security vulnerabilities, threats, and configuration weaknesses
SEO Crawler
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
Uncategorized
UncategorizedNot yet assigned a type
Undocumented AI Agent
Undocumented AI AgentCrawls websites without disclosing its purpose, collecting data for an unknown AI use case
Hover over each agent type for more information about what they do
Agent types with the most activity
bingbot
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
9.1%
AhrefsBot
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
7.3%
Googlebot
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
6.5%
Known Agent
DEV
Developer HelperAssists with testing, debugging, and ensuring website functionality
6.5%
ClaudeBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
3.4%
PetalBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
3.3%
SemrushBot
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
2.8%
facebookexternalhit
FTCH
FetcherRetrieves web page metadata to power app features like link previews or feeds
2.7%
Baiduspider
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
2.5%
Amazonbot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
2.4%
meta-externalagent
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
2.3%
Amzn-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
2.2%
ChatGPT-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
2.1%
MJ12bot
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
2.0%
DotBot
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
1.9%
Agents with the most activity
Operators with the most activity
These bots scrape website content to train AI models. Some belong to AI companies, while others belong to third-party services that resell the data. Automatic Robots.txt can block unwanted scraping. Included agent types include AI Data Providers and AI Data Scrapers.
AI Data Provider
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
AI Data Scraper
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
AI training activity by agent type over time
Computers and Electronics
6.3%
Business and Industrial
5.9%
Internet and Telecom
5.7%
Website categories with most activity
ClaudeBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
25.9%
Amazonbot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
17.9%
meta-externalagent
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
17.6%
GPTBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
7.1%
ShapBot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
7.0%
Bytespider
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
6.6%
Reflectionbot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
3.8%
YouBot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
2.7%
GoogleOther
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
2.6%
CCBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
2.3%
Timpibot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
1.3%
DeepSeekBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
1.2%
Diffbot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
1.0%
VelenPublicWebCrawler
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
0.8%
FacebookBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
0.5%
Agents doing the most AI training
Operators doing the most AI training
These bots fetch website content in real time to power AI assistants, coding agents, and other retrieval-augmented generation (RAG) tasks. Pages inform responses on the spot, such as when an assistant summarizes an article or a coding agent references documentation. Included agent types include AI Assistants and AI Coding Agents.
AI Assistant
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
AI Coding Agent
AI Coding AgentFetches documentation and other resources to help build software
AI fetching activity by agent type over time
Computers and Electronics
2.0%
Travel and Transportation
1.8%
Internet and Telecom
1.8%
Business and Industrial
1.5%
Website categories with most activity
ChatGPT-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
86.9%
Claude-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
4.4%
DuckAssistBot
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
4.0%
Claude-Code
CODE
AI Coding AgentFetches documentation and other resources to help build software
2.2%
Shap-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.7%
Cursor
CODE
AI Coding AgentFetches documentation and other resources to help build software
0.6%
Google-NotebookLM
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.4%
opencode
CODE
AI Coding AgentFetches documentation and other resources to help build software
0.3%
Perplexity-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.2%
MistralAI-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.1%
GoogleAgent-URLContext
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.1%
Gemini-Deep-Research
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
Code
CODE
AI Coding AgentFetches documentation and other resources to help build software
0.0%
amazon-QBusiness
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
Amzn-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.0%
Agents doing the most AI fetching
Operators doing the most AI fetching
These bots crawl website content so it can be surfaced in AI search engines and AI-generated answers. Those answers often include citations or links back to the source pages. Included agent types include AI Search Crawlers.
AI Search Crawler
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
AI search indexing activity by agent type over time
Travel and Transportation
7.2%
Books and Literature
5.5%
Arts and Entertainment
5.0%
Computers and Electronics
4.6%
Website categories with most activity
PetalBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
26.3%
Amzn-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
17.3%
Claude-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
13.0%
Applebot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
12.9%
meta-webindexer
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
12.1%
OAI-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
10.5%
LinkupBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
3.2%
PerplexityBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
2.5%
ExaSearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
1.0%
AIWebIndex
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.4%
xAI-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.3%
AzureAI-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.3%
Google-CloudVertexBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.3%
Anomura
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.0%
AddSearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.0%
Agents doing the most AI search indexing
Operators doing the most AI search indexing
These bots use browsers to autonomously navigate websites, click through pages, and make decisions to complete tasks for people. Agentic UX best practices and Google PageSpeed Insights help evaluate how well websites support them. Included agent types include AI Agents.
↑ 45%
Compared to the previous 90 days
The average duration of a session
↑ 103%
Compared to the previous 90 days
The average number of pages visited per session
AI Agent
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
AI browsing activity by agent type over time
Computers and Electronics
0.1%
Business and Industrial
0.0%
Arts and Entertainment
0.0%
Internet and Telecom
0.0%
Travel and Transportation
0.0%
Books and Literature
0.0%
Website categories with most activity
Google-Agent
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
36.5%
ChatGPT Agent
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
36.4%
Manus-User
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
27.1%
NovaAct
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
GoogleAgent-Mariner
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
AmazonBuyForMe
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
TwinAgent
AGNT
AI AgentUses an actual web browser to autonomously complete complex tasks on behalf of a human user
0.0%
Agents doing the most AI browsing
Operators doing the most AI browsing
See which robots.txt rules are set across the web and how well agents follow them. An agent's Robots.txt Effectiveness measures the effectiveness of a disallow rule for it by estimating how much the agent reduces its traffic after it's blocked.
How often
all agents across all agent types follow robots.txt rules
proximic
INT
Intelligence GathererAnalyzes web content for brand safety, competitive insights, and ad targeting
100.0%
Googlebot
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
100.0%
AhrefsSiteAudit
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
99.9%
panscient.com
INT
Intelligence GathererAnalyzes web content for brand safety, competitive insights, and ad targeting
99.8%
YouBot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
99.7%
LinkupBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
99.6%
Barkrowler
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
99.6%
Claude-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
99.6%
ExaSearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
99.2%
GPTBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
99.1%
meta-externalagent
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
99.1%
VelenPublicWebCrawler
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
99.0%
Amazonbot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
98.8%
PerplexityBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
98.8%
SERankingBacklinksBot
SEO
SEO CrawlerAnalyzes website structure and content to identify SEO improvement opportunities
98.8%
Agents with the best Robots.txt Effectiveness percentages
Baiduspider
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
64.7%
ShapBot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
75.5%
QwenBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
79.0%
MoonshotBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
79.8%
DeepSeekBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
82.5%
ChatGLM-Spider
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
83.4%
PanguBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
84.7%
cohere-ai
UND
Undocumented AI AgentCrawls websites without disclosing its purpose, collecting data for an unknown AI use case
86.1%
Bravebot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
87.4%
OAI-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
88.4%
Sogou web spider
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
88.4%
Applebot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
92.0%
PetalBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
92.9%
CCBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
93.8%
ChatGPT-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
94.2%
Agents with the worst Robots.txt Effectiveness percentages
●
CCBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
24.6%
●
Bytespider
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
22.5%
●
GPTBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
21.5%
●
ClaudeBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
21.4%
●
omgili
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
19.8%
●
Diffbot
PVDR
AI Data ProviderCrawls websites to supply structured content to AI systems as a third-party service
19.5%
●
anthropic-ai
UND
Undocumented AI AgentCrawls websites without disclosing its purpose, collecting data for an unknown AI use case
19.2%
●
meta-externalagent
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
19.1%
●
cohere-ai
UND
Undocumented AI AgentCrawls websites without disclosing its purpose, collecting data for an unknown AI use case
18.9%
●
PerplexityBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
18.1%
●
Claude-Web
UND
Undocumented AI AgentCrawls websites without disclosing its purpose, collecting data for an unknown AI use case
17.6%
●
Omgilibot
UNC
UncategorizedNot yet assigned a type
17.0%
●
FacebookBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
16.9%
Agents blocked by the most top websites
See which agents are most frequently impersonated, and how spoofing activity changes over time. A visit is considered spoofed when it claims a recognized agent identity but fails that agent's supported authentication method, such as verified IP or HTTP message signature (Web Bot Auth).
Active Threat: AI Bot Spoofing Campaign
We are observing a widespread campaign impersonating AI bots to scan websites for vulnerabilities. The attacker appears to be targeting credential and configuration paths used by AI coding tools.
Contact us for more information, or
inspect your own traffic.
The percentage of impersonated website traffic for each agent identity over time
●
Googlebot
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
0.4%
●
ChatGPT-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.2%
●
OAI-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.1%
●
PerplexityBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.1%
●
ClaudeBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
0.1%
●
GPTBot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
0.1%
●
Applebot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.1%
●
Claude-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.1%
●
Perplexity-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.1%
●
Amzn-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.1%
●
MistralAI-User
ASST
AI AssistantFetches website content in response to a user prompt, to include in an AI-generated answer
0.1%
●
bingbot
SRCH
Search Engine CrawlerSystematically scans and indexes web pages to include in search results
0.0%
●
Amazonbot
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
0.0%
●
Claude-SearchBot
SRCH
AI Search CrawlerIndexes website content to possibly include as citations in AI-powered search results
0.0%
●
GoogleOther
SCRP
AI Data ScraperDownloads website content to include in datasets used for training AI models such as LLMs
0.0%
The most impersonated agent identities
/.config/anthropic/credentials/default.json
/.claude/settings.json
/.claude.json
/.hermes/.env
/.openclaw/.env
/.codex/config.toml
/.continue/config.json
/.aider.conf.yml
/service-account.json
/serviceaccountkey.json
/service_account.json
/firebase-adminsdk.json
/firebase-service-account.json
/.aws/credentials
/.aws/config
/.s3cfg
/.boto
/.npmrc
/.env.example
/.env.local
/.env.production
/.env.backup
/.env.old
/backend/.env
/api/.env
/admin/.env
/dockerfile
/docker-compose.yaml
/.docker/config.json
/terraform.tfstate
/credentials.json
/secrets.json
/secrets.yml
/key.json
/rclone.conf
Examples of recent top targeted request paths
See which AI platforms like ChatGPT, Perplexity, and Gemini cite websites and send them human referral traffic. Citations are estimated. Google's guide explains how websites can optimize their content to be more visible in AI chat responses (GEO).
Computers and Electronics
1.9%
Travel and Transportation
1.8%
Internet and Telecom
1.7%
Business and Industrial
1.5%
Website categories most frequently cited in AI chat responses
Travel and Transportation
0.3%
Computers and Electronics
0.2%
Business and Industrial
0.1%
Internet and Telecom
0.0%
Website categories receiving the most referrals from AI chat
The Index updates daily using completed days of traffic, security, and referral data from more than 5,000 websites using Agent Analytics and AI Chat Referral Tracking. Percentage changes compare equal-length periods. Agent data comes from the Agent Directory, and website categories follow Google AdSense. Participating websites are not a random sample, and the qualifying set changes over time, so results show observed directional trends rather than a census of global web traffic.
Only website-days meeting minimum activity and data-quality requirements are included; internal, test, incomplete, and anomalous data is excluded. Bot percentages use all visits, while AI chat referral percentages use estimated human visits after identified bots are removed. Rates are calculated per website-day and then averaged, giving each website equal weight regardless of traffic. Daily charts are not smoothed in order to preserve natural seasonality.
Robots.txt Effectiveness estimates the traffic reduction associated with blocking an agent sitewide using Disallow: /. For each agent and day, the baseline allowed rate includes every qualifying website that allows it, even those with zero agent visits:
For each blocked website, expected visits are:
Given observed disallowed visits O, its score is:
A 0% score means visits met or exceeded expectations, 50% means half as many visits as expected, and 100% means none were observed.
Scores require meaningful activity across multiple allowed and blocked websites over at least 7 of the latest 14 completed days. Each qualifying website-day has equal weight, and the headline score gives each qualifying agent equal weight. Only verified visits count for agents supporting authentication. This observational score measures outcomes, not intent or causation.
Top Blocked Bots is calculated separately from robots.txt scans of Similarweb's top 1,000 websites.
Spoofing statistics count visits that claim a known agent's identity but fail a supported authentication method, such as IP verification or HTTP message signatures (Web Bot Auth). Each agent's daily rate is calculated against total visits per website and then averaged across websites. A failure suggests impersonation but does not identify the actual software or operator. Agents without supported authentication are excluded.
AI chat referrals count observed human visits with a recognized AI platform in the referrer or campaign source; visits without that data cannot be attributed. Citation estimates use requests from agents that retrieve content for AI platforms. Those requests may inform a response but do not prove a user saw a citation, so citation results are directional rather than exact counts.
Can journalists and media organizations use this data?
Absolutely. You may cite The Agentic Web Index with attribution and a link to this page. For interviews, fact-checking, background context, or a more specific breakdown for a story, contact us and include your deadline.
Do you work with researchers?
Absolutely. We welcome thoughtful research into how agents and bots are changing the web. Tell us about your research question, timeframe, and intended use. Depending on the scope and data constraints, we may be able to provide additional context, compare approaches, or explore a joint analysis.
Can I request a specific analysis?
Yes. If you need a breakdown by agent, operator, activity type, website category, or time period that is not shown here, contact us. When the underlying data supports it, we can examine the question and provide a focused analysis.
How do I see these trends on my own website?