What is VelenPublicWebCrawler?

VelenPublicWebCrawler is a web crawler developed by Velen for Hunter that analyzes millions of publicly accessible internet pages every month. The bot builds business datasets and machine learning models while crawling respectfully with a minimum 2-second delay between requests. Agent Analytics can track when it visits your website.

Overview

Operated By Velen
Source Official Website
Expected To Follow Robots.txt Yes
Insights Last Updated July 31, 2026

Do you operate this agent? Contact us to suggest an update.

Category

AI Data Scraper
Downloads website content to include in datasets used for training AI models such as LLMs

Expected Behavior

VelenPublicWebCrawler crawls websites to build AI training datasets, and its operator decides which sites, how often, and how deep it goes. Expect broad sweeps that fetch far more pages per visit than a search crawler would, on an unpredictable schedule. Volume can run heavy while a collection pass is underway, then stop entirely.

VelenPublicWebCrawler's User Agent

User Agent Mozilla/5.0 (compatible; VelenPublicWebCrawler/1.0; +https://velen.io)

How To Block VelenPublicWebCrawler With Robots.txt

Add this rule to your robots.txt file to block VelenPublicWebCrawler from accessing your entire website, or use Automatic Robots.txt to block all AI data scrapers at once. You can customize which pages are blocked by swapping out / for a different path.

User-agent: VelenPublicWebCrawler # https://knownagents.com/agents/velenpublicwebcrawler
Disallow: /

Global Insights for VelenPublicWebCrawler

As of July 31, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the world's top 1000 websites and their robots.txt files.

Robots.txt Blocked Percentage

5%
5% of top websites are blocking VelenPublicWebCrawler
Learn How →

Country of Origin

Belgium
VelenPublicWebCrawler normally visits From Belgium

Robots.txt Blocking Trend

5% of top websites block VelenPublicWebCrawler in their robots.txt files.

Overall AI Data Scraper Traffic

1.7% of all web traffic came from AI data scrapers.

Top Visited Website Categories

Law and Government
Science
Jobs and Education
Real Estate
Home and Garden

The types of websites most frequently visited by VelenPublicWebCrawler.

Frequently Asked Questions

Should I Block VelenPublicWebCrawler?

Block it if you want control over how your work trains AI models. VelenPublicWebCrawler downloads your content into training datasets without attribution, compensation, or any promise of traffic back. The case for allowing it is reach, because models trained on your content can surface your brand in AI answers. Neither choice changes your Google rankings. For comparison, 5% of the top websites we track already have robots.txt rules for VelenPublicWebCrawler.


Does VelenPublicWebCrawler Follow Robots.txt Rules?

Yes. VelenPublicWebCrawler is expected to follow robots.txt rules, so a disallow rule is the right first move. Automatic Robots.txt adds and maintains that rule for you, and Agent Analytics confirms VelenPublicWebCrawler actually follows it.


Does VelenPublicWebCrawler Access Private Content?

VelenPublicWebCrawler targets public content, but the boundary is not always respected. Some training data scrapers go after paywalled or gated pages when the operator wants that data. If a page loads without signing in, assume it can be collected.


Why Is VelenPublicWebCrawler Visiting My Website?

Your content matched what Velen wants in a training dataset. VelenPublicWebCrawler discovers pages mechanically, through links from other sites, sitemaps, and seed lists, not because anyone chose your site personally.


How Can I Tell if VelenPublicWebCrawler Is Visiting My Website?

Agent Analytics tracks VelenPublicWebCrawler visits in real time alongside every other known AI agent, crawler, and scraper. You can also check your server logs for requests whose user agent string contains "VelenPublicWebCrawler". Look for broad sweeps that fetch large numbers of pages in sequence. Keep in mind that VelenPublicWebCrawler doesn't publish a verification method, so any client can claim its user agent string and a log match is a hint rather than proof.