What is ApifyWebsiteContentCrawler?

ApifyWebsiteContentCrawler is a web crawler by Apify that extracts and downloads full website content for use in AI, data analysis, and automation workflows. You can set up Agent Analytics to see when ApifyWebsiteContentCrawler visits your website.

Overview

Operated By Apify
Source Official Website
Expected To Follow Robots.txt Yes
Insights Last Updated September 14, 2026

Do you operate this agent? Contact us to suggest an update.

Category

AI Data Provider
Crawls websites to supply structured content to AI systems as a third-party service

Expected Behavior

ApifyWebsiteContentCrawler supports several kinds of automated activity, so its traffic pattern depends on the task. It may perform broad crawls to collect content or build an index, make targeted requests for specific pages, or shift between these behaviors as demand changes.

ApifyWebsiteContentCrawler's User Agent

User Agent apifywebsitecontentcrawler

How To Block ApifyWebsiteContentCrawler With Robots.txt

Add this rule to your robots.txt file to block ApifyWebsiteContentCrawler from accessing your website, or use Automatic Robots.txt to block all AI data providers at once. You can customize which pages are blocked by swapping out / for a different path.

User-agent: ApifyWebsiteContentCrawler # https://knownagents.com/agents/apifywebsitecontentcrawler
Disallow: /

Insights for ApifyWebsiteContentCrawler

As of September 14, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the top 1000 websites and their robots.txt files.

Robots.txt Blocked Percentage

8%
8% of top websites are blocking ApifyWebsiteContentCrawler
Learn How →

Country of Origin

United States
ApifyWebsiteContentCrawler normally visits From the United States

Robots.txt Blocking Trend

8% of top websites block ApifyWebsiteContentCrawler in their robots.txts.

Overall AI Data Provider Traffic

0.2% of all web traffic came from AI data providers.

Frequently Asked Questions

Should I Block ApifyWebsiteContentCrawler?

It depends on which uses you want to support. ApifyWebsiteContentCrawler may supply your content to enterprises, including AI companies, or use it in consumer-facing products. Those uses can include AI training, search, and retrieval. Allowing it may expand your visibility across those products and services, while blocking it limits that reach without affecting traditional search rankings. For context, 8% of the top websites we track currently have robots.txt rules for ApifyWebsiteContentCrawler.


Does ApifyWebsiteContentCrawler Follow Robots.txt Rules?

Yes. ApifyWebsiteContentCrawler is expected to follow robots.txt rules, so a disallow rule is the appropriate first step. Automatic Robots.txt can add and maintain that rule, while Agent Analytics lets you verify whether ApifyWebsiteContentCrawler respects it.


Does ApifyWebsiteContentCrawler Access Private Content?

No special access. ApifyWebsiteContentCrawler can reach public pages, and some providers use proxy networks that can bypass rate limits or geographic restrictions. Protect sensitive content with authentication rather than relying on those softer barriers.


How Can I Tell if ApifyWebsiteContentCrawler Is Visiting My Website?

Agent Analytics tracks visits from ApifyWebsiteContentCrawler and every other known AI agent, crawler, and scraper, then authenticates each visit using ApifyWebsiteContentCrawler's published verification method. You can also check your server logs for requests whose user-agent string contains "ApifyWebsiteContentCrawler". Traffic may range from long crawl sequences across your site to isolated requests for specific pages, depending on whether it is collecting, indexing, or fetching content on demand. A matching log entry alone is not proof because any bot can claim to be ApifyWebsiteContentCrawler.


Why Is ApifyWebsiteContentCrawler Visiting My Website?

One of Apify's customers may have requested data from your site, or your pages may be part of ApifyWebsiteContentCrawler's standing index. A single fetch can support multiple downstream applications.