# Controlling Web Crawlers For a Site

**URL:** <https://meta.discourse.org/t/controlling-web-crawlers-for-a-site/244925>\
**Category:** Site Management\
**Tags:** how-to\
**Created:** [November 9, 2022, 12:04pm UTC](https://meta.discourse.org/t/controlling-web-crawlers-for-a-site/244925 "2022-11-09T12:04:55Z")\
**Posts on this page:** 1\
**Showing post:** 1

<div class="post-metadata">

**Author:** ![Discourse](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/discourse/32/148734_2.png) [@Discourse](https://meta.discourse.org/u/Discourse)\
**Post date:** [November 9, 2022, 12:04pm UTC](https://meta.discourse.org/t/controlling-web-crawlers-for-a-site/244925/1 "2022-11-09T12:04:55Z")

</div>

> 🔖 This guide explains how to manage web crawlers on your Discourse site.
> 
> 🙋 Required user level: Administrator

Web crawlers can significantly impact your site’s performance by increasing pageviews and server load.

When a site notices a spike in their pageviews it’s important to check how web crawlers fit into the mix.

* * *

## Checking for crawler activity

To see if crawlers are affecting your site, navigate to the **Site Traffic** report (`/admin/reports/site_traffic` ) from your admin dashboard. This report breaks down pageview numbers from logged-in browser users, anonymous browser users, crawlers, and other sources.

_A site where crawlers work normally:_

 ![image](https://global.discourse-cdn.com/meta/original/4X/e/8/9/e89dc5d1f634957dcbed7f5a821ed710ad260401.png)

_A site where crawlers are out of control:_

 ![image](https://global.discourse-cdn.com/meta/original/4X/e/2/e/e2eb7e8b84bcd7a16a76fe623e3ccc3efa2cb15b.png)

## Identifying specific crawlers

Go to the **Web Crawler User Agent** report (`/admin/reports/web_crawlers`) to find a list of web crawler names sorted by pageview count.

When a problematic web crawler hits the site, the number of its pageviews will be much higher than the other web crawlers. Note that there may be a number of malicious web crawlers at work at the same time.

## Blocking and limiting crawlers

It is a good habit not to block the crawlers of the main search engines, such as [Google](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers), [Bing](https://www.bing.com/webmasters/help/which-crawlers-does-bing-use-8c184ec0), [Baidu](https://developers.whatismybrowser.com/useragents/explore/software_name/baidu-spider/) (Chinese), [Yandex](https://yandex.com/support/webmaster/robot-workings/check-yandex-robots.html) (Russian), [Naver](https://searchadvisor.naver.com/guide/seo-basic-firewall) (Korean), [DuckDuckGo](https://developers.whatismybrowser.com/useragents/explore/software_name/duckduckgo/), [Yahoo](https://help.yahoo.com/kb/SLN22600.html) and others, based on your country.

When a web crawler is out of control there is a good chance that the same crawler has hit other sites and someone else has already asked for information or created reports about it that will be useful to understand whether to limit or block that particular crawler.

Note that some crawlers may contribute a large number of pageviews if you use third-party services to monitor or add functionality to your site via scripts, etc.

To obtain a record of untrustworthy web crawlers, you may refer to this list, [https://github.com/mitchellkrogza/apache-ultimate-bad-bot-blocker/blob/master/robots.txt/robots.txt](https://github.com/mitchellkrogza/apache-ultimate-bad-bot-blocker/blob/master/robots.txt/robots.txt)

## Adjusting crawler settings

Under **Admin \> Settings** there are some settings that can help rate limit specific crawlers:

- **Slow down crawlers** using:

- **Block crawlers** with:

- **Allow only specific crawlers** with:

Ensure you know the accurate user agent name for the crawlers you wish to control. If you adjust any of the settings above and do not see a reduction in pageviews of that agent, you may want to double check that you are using the proper name.

When in doubt about how to act, always start with the “slow down” option rather than a full block. Check over time if there are improvements. You can proceed with a full block if you do not notice appreciable results.

> Last edited by @SaraDev 2024-09-11T19:32:59Z
> 
> > **Check document**
> >
> > Perform check on document:

---

_[View the full topic](https://meta.discourse.org/t/controlling-web-crawlers-for-a-site/244925)._
