Operated by Allen Institute for AI
A specific AI2 crawler used to build the Dolma dataset, a massive open corpus for training language models.
A specific AI2 crawler used to build the Dolma dataset, a massive open corpus for training language models.
Ai2Bot-Dolma is an AI data-collection crawler operated by Allen Institute for AI. It harvests web content to build or expand training datasets for large language models (LLMs). Unlike search crawlers, Ai2Bot-Dolma does NOT influence your page ranking in any search engine. The user-agent string Ai2Bot-Dolma can be safely blocked via robots.txt, meta tags (noai), or the emerging llms.txt standard without any SEO penalty. Robots.txt is voluntary; for hard enforcement, combine it with server-level IP blocking.
User-agent: Ai2Bot-Dolma / Disallow: / without any SEO penalty. This is the recommended approach if you want to opt out of Allen Institute for AI's LLM training datasets.<code>User-agent: Ai2Bot-Dolma</code> — Matching is case-insensitive. Robots.txt is fetched from the root of each subdomain separately.
Understanding Ai2Bot-Dolma's purpose helps you decide whether to allow or block it.
Ai2Bot-Dolma. This is the exact string you must use in robots.txt, Nginx, Apache, or Cloudflare firewall rules to target this bot. User-agent matching in robots.txt is case-insensitive, but the string must be spelled correctly. You can verify that a request genuinely comes from Ai2Bot-Dolma by performing a reverse-DNS lookup on the source IP — legitimate bots resolve back to their operator's domain.User-agent: Ai2Bot-Dolma / Disallow: / without any SEO penalty. This is the recommended approach if you want to opt out of Allen Institute for AI's LLM training datasets./robots.txt file:
User-agent: Ai2Bot-Dolma Disallow: /This instructs Ai2Bot-Dolma not to crawl any path on your site. The Disallow: / directive covers the entire domain including subfolders. To only block specific sections, replace / with the path (e.g.,
Disallow: /blog/). Note: robots.txt is publicly readable — any bot or human can inspect it at yourdomain.com/robots.txt.Ai2Bot-Dolma (case-insensitive grep: grep -i "Ai2Bot-Dolma" /var/log/nginx/access.log). You can also check Google Search Console → Coverage → Crawl Stats for Googlebot variants. For Ai2Bot-Dolma specifically, filter by user-agent in your log analysis tool (GoAccess, AWStats, etc.).Disallow: / you can restrict Ai2Bot-Dolma to specific paths:
User-agent: Ai2Bot-Dolma Disallow: /private/ Disallow: /staging/ Allow: /This allows Ai2Bot-Dolma everywhere except the listed paths. Path matching in robots.txt uses prefix matching —
Disallow: /private/ blocks /private/page.html but NOT /public/private/.<meta name="Ai2Bot-Dolma" content="noai, noimageai, noindex"> to your pages.
2. Add a llms.txt file at your domain root (emerging standard).
3. Use Cloudflare WAF or Nginx to return 403 for this user-agent.
4. Consider IP blocklists for Allen Institute for AI's known crawler IP ranges.<meta name="Ai2Bot-Dolma" content="noindex">
• **X-Robots-Tag HTTP header**: X-Robots-Tag: noai, noimageai
• **llms.txt**: Add a /llms.txt file (similar to robots.txt but for LLMs)
• **Server block**: Return 403 or 429 for this user-agent via WAF or Nginx
Using multiple layers provides the strongest protection.Check instantly with our free AI Bot Checker
Check Your Website