What is ClueWeb-Crawler?

ClueWeb-Crawler is an uncategorized agent. Agent Analytics can track when it visits your website.

Overview

Expected To Follow Robots.txt Yes
Insights Last Updated July 17, 2026

Category

Uncategorized
Not yet assigned a type

Expected Behavior

ClueWeb-Crawler has no established behavior profile yet. Read its pattern in Agent Analytics: steady polite crawling suggests an unannounced index, while fast deep sweeps suggest scraping. Its user agent string and IP addresses are the best clues to who runs it.

ClueWeb-Crawler's User Agent

User Agent ClueWeb-Crawler/1.0 (+https://boston.lti.cs.cmu.edu/CMU-ClueWeb-Crawler/; mailto:cmu-clueweb-crawler@andrew.cmu.edu)

How To Block ClueWeb-Crawler With Robots.txt

Add this rule to your robots.txt file to block ClueWeb-Crawler from accessing your entire website, or use Automatic Robots.txt to block all uncategorized agents at once. You can customize which pages are blocked by swapping out / for a different path.

User-agent: ClueWeb-Crawler # https://knownagents.com/agents/clueweb-crawler
Disallow: /

Global Insights for ClueWeb-Crawler

As of July 17, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the world's top 1000 websites and their robots.txt files.

Robots.txt Blocked Percentage

0%
0% of top websites are blocking ClueWeb-Crawler
Learn How →

Country of Origin

United States
ClueWeb-Crawler normally visits From the United States

Robots.txt Blocking Trend

0% of top websites block ClueWeb-Crawler in their robots.txt files.

Overall Uncategorized Traffic

1.9% of all web traffic came from uncategorized agents.

Top Visited Website Categories

Computers and Electronics
Hobbies and Leisure
Internet and Telecom
Food and Drink
Finance

The types of websites most frequently visited by ClueWeb-Crawler.

Frequently Asked Questions

Should I Block ClueWeb-Crawler?

Watch it first. ClueWeb-Crawler has not been categorized yet, so there is no track record to lean on. Check which pages it requests and how often, and block it if the traffic is heavy or pointed at content you would not give an unknown bot. Almost none of the top websites we track have robots.txt rules for ClueWeb-Crawler right now.


Does ClueWeb-Crawler Follow Robots.txt Rules?

Yes. ClueWeb-Crawler is expected to follow robots.txt rules, so a disallow rule is the right first move. Automatic Robots.txt adds and maintains that rule for you, and Agent Analytics confirms ClueWeb-Crawler actually follows it.


Does ClueWeb-Crawler Access Private Content?

Unknown. ClueWeb-Crawler has no documented scope, so watch what it actually requests. Repeated hits on login or admin paths are the signal to block it.


Why Is ClueWeb-Crawler Visiting My Website?

Nobody knows yet. ClueWeb-Crawler is undocumented, so until its operator explains it, its behavior on your site is the only evidence of its intent.


How Can I Tell if ClueWeb-Crawler Is Visiting My Website?

Agent Analytics tracks ClueWeb-Crawler visits in real time alongside every other known AI agent, crawler, and scraper. You can also check your server logs for requests whose user agent string contains "ClueWeb-Crawler". Log the paths it requests and how often, since nothing about it is documented. Keep in mind that ClueWeb-Crawler doesn't publish a verification method, so any client can claim its user agent string and a log match is a hint rather than proof.