What is cohere-training-data-crawler?
cohere-training-data-crawler is a web crawler operated by Cohere to download training data for its LLMs (Large Language Models) that power its enterprise AI products. You can set up Agent Analytics to see when cohere-training-data-crawler visits your website.
Overview
| Operated By | Cohere |
| Expected To Follow Robots.txt | Yes |
| Insights Last Updated | September 15, 2026 |
Do you operate this agent? Contact us to suggest an update.
Category
Expected Behavior
cohere-training-data-crawler tends to make broad, high-volume sweeps across many pages. Timing is unpredictable: traffic may remain heavy throughout a collection pass, then stop entirely.
cohere-training-data-crawler's User Agent
| User Agent | cohere-training-data-crawler |
How To Block cohere-training-data-crawler With Robots.txt
Add this rule to your robots.txt file to block cohere-training-data-crawler from accessing your website, or use Automatic Robots.txt to block all AI data scrapers at once. You can customize which pages are blocked by swapping out / for a different path.
User-agent: cohere-training-data-crawler # https://knownagents.com/agents/cohere-training-data-crawler
Disallow: /
Insights for cohere-training-data-crawler
As of September 15, 2026, this data reflects agent visits measured across thousands of websites using Agent Analytics, combined with daily scans of the top 1000 websites and their robots.txt files.
Robots.txt Blocked Percentage
Country of Origin
Robots.txt Blocking Trend
13% of top websites block cohere-training-data-crawler in their robots.txts.
Overall AI Data Scraper Traffic
2.7% of all web traffic came from AI data scrapers.
Frequently Asked Questions
Should I Block cohere-training-data-crawler?
Block cohere-training-data-crawler if you want more control over whether your work is used for AI training. Allowing it may increase the chance that your brand appears in AI-generated answers, but either choice leaves traditional search rankings unchanged. For context, 13% of the top websites we track currently have robots.txt rules for cohere-training-data-crawler.
Does cohere-training-data-crawler Follow Robots.txt Rules?
Yes. cohere-training-data-crawler is expected to follow robots.txt rules, so a disallow rule is the appropriate first step. Automatic Robots.txt can add and maintain that rule, while Agent Analytics lets you verify whether cohere-training-data-crawler respects it.
Does cohere-training-data-crawler Access Private Content?
Not through legitimate access. cohere-training-data-crawler primarily targets public content, but some training-data scrapers also attempt to collect gated or paywalled pages. If content loads without authentication, assume it can be collected.
How Can I Tell if cohere-training-data-crawler Is Visiting My Website?
Agent Analytics tracks cohere-training-data-crawler visits in real time. You can also check your server logs for requests whose user-agent string contains "cohere-training-data-crawler". Look for high page counts, short gaps between requests, and deep traversal through linked content. Because cohere-training-data-crawler does not publish a verification method, any client can claim its identity and a log match is only a clue.
Why Is cohere-training-data-crawler Visiting My Website?
Your content matched the criteria Cohere set for a training dataset. cohere-training-data-crawler typically discovers pages through external links, sitemaps, and seed lists rather than because someone selected your site individually.