What is Nutch?

Nutch is an open-source web crawler developed by the Apache Software Foundation. It is commonly used for large-scale web scraping. Agent Analytics can track when it visits your website.

Overview

Expected To Follow Robots.txt Yes
Insights Last Updated July 12, 2026

Category

Scraper
Extracts large amounts of web data, often without explicit website permission

Expected Behavior

Nutch behaves however its operator configured it, and scrapers as a category are the least polite bots on the web. Expect anything from slow careful extraction to rapid page hammering, robots.txt ignored, and user agent strings that change when blocked. Watch its volume and speed rather than trusting its label.

Nutch's User Agent

User Agent MaxPointCrawler/Nutch-1.19 (valassis.crawler at valassis dot com)

How To Block Nutch With Robots.txt

Add this rule to your robots.txt file to block Nutch from accessing your entire website, or use Automatic Robots.txt to block all scrapers at once. You can customize which pages are blocked by swapping out / for a different path.

User-agent: Nutch # https://knownagents.com/agents/nutch
Disallow: /

Nutch Global Insights

As of July 12, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the world's top 1000 websites and their robots.txt files.

Robots.txt Blocked Percentage

3%
3% of top websites are blocking Nutch
Learn How →

Country of Origin

United States
Nutch normally visits From the United States

Robots.txt Blocking Trend

3% of top websites block Nutch in their robots.txt files.

Overall Scraper Traffic

0.0% of all web traffic came from scrapers.

Top Visited Website Categories

Reference
Sports
Food and Drink
News
Science

The types of websites most frequently visited by Nutch.

Frequently Asked Questions

Should I Block Nutch?

Often yes. Nutch extracts content at scale, and scrapers in this category commonly republish or resell what they take. There is rarely an upside for the website being scraped, and copies of your content elsewhere can compete with your own pages in search. For comparison, 3% of the top websites we track already have robots.txt rules for Nutch.


Does Nutch Follow Robots.txt Rules?

Yes. Nutch is expected to follow robots.txt rules, so a disallow rule is the right first move. Automatic Robots.txt adds and maintains that rule for you, and Agent Analytics confirms Nutch actually follows it.


Does Nutch Access Private Content?

Assume Nutch will try. Scrapers routinely ignore robots.txt, and some go after paywalled or gated content when it has value. Real authentication stops most of them. Politeness conventions stop almost none of them.


Why Is Nutch Visiting My Website?

Your site has data Nutch's operator wants, like prices, listings, contact details, or articles. Scrapers target sites deliberately, so repeated visits mean your content specifically is the goal.


How Can I Tell if Nutch Is Visiting My Website?

Agent Analytics tracks Nutch visits in real time alongside every other known AI agent, crawler, and scraper. You can also check your server logs for requests whose user agent string contains "Nutch". Look for fast sequential requests across many pages. Keep in mind that Nutch doesn't publish a verification method, so any client can claim its user agent string and a log match is a hint rather than proof.

Sources