# 爬虫数量是否过高？

**URL:** https://meta.discourse.org/t/crawlers-very-high/151339
**Category:** Support
**Created:** [2020年五月13日 15:00 UTC](https://meta.discourse.org/t/crawlers-very-high/151339 "2020-05-13T15:00:15Z")
**Posts on this page:** 7
**Page:** 1

<div class="post-metadata">

### Author: ![Bank\_Live](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/bank_live/32/75893_2.png) [@Bank\_Live](https://meta.discourse.org/u/Bank_Live)
#### Post date: [2020年五月13日 15:00 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/1 "2020-05-13T15:00:15Z")

</div>

我认为这个值太高了。有什么问题吗？

 ![图片](https://global.discourse-cdn.com/meta/original/3X/e/b/eb368d8413369b9344bd82f35fb55c241a5fc904.jpeg)

---

<div class="post-metadata">

### Author: ![Stranik](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/stranik/32/85638_2.png) [@Stranik](https://meta.discourse.org/u/Stranik)
#### Post date: [2020年五月13日 15:36 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/2 "2020-05-13T15:36:48Z")

</div>

您可以查看具体是哪些爬虫：

```plaintext
/admin/reports/web_crawlers

```

并将它们添加到此页面：

```plaintext
admin/site_settings/category/all_results?filter=crawler user agents

```

> [@How to block all crawlers but Google's](https://meta.discourse.org/t/how-to-block-all-crawlers-but-googles/62431):
>
> About 1/3 of our traffic is from crawlers (about 250K last month). Is there a way to block these but allow Google’s crawlers?

---

<div class="post-metadata">

### Author: ![Bank\_Live](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/bank_live/32/75893_2.png) [@Bank\_Live](https://meta.discourse.org/u/Bank_Live)
#### Post date: [2020年五月13日 15:58 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/3 "2020-05-13T15:58:30Z")

</div>

> [@Stranik](#):
>
> `/admin/reports/web_crawlers`

 ![图片](https://global.discourse-cdn.com/meta/original/3X/d/f/dfbd9aa84cdb2bbe5f33d6ebc5839648e7438972.png)

---

<div class="post-metadata">

### Author: ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)
#### Post date: [2020年五月13日 15:58 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/4 "2020-05-13T15:58:34Z")

</div>

最近由于 Google 更改了 robots.txt 的行为，关于爬取规则出现了很多变动，因此这_可能_是正常的。

---

<div class="post-metadata">

### Author: ![Bank\_Live](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/bank_live/32/75893_2.png) [@Bank\_Live](https://meta.discourse.org/u/Bank_Live)
#### Post date: [2020年五月13日 16:00 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/5 "2020-05-13T16:00:23Z")

</div>

好的，非常感谢。

---

<div class="post-metadata">

### Author: ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)
#### Post date: [2020年五月13日 16:02 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/6 "2020-05-13T16:02:46Z")

</div>

除非您的数据显示并非如此，否则存在一个恶意爬虫！请通过您的站点设置将其屏蔽。

---

<div class="post-metadata">

### Author: ![system](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/system/32/443519_2.png) [@system](https://meta.discourse.org/u/system)
#### Post date: [2020年六月12日 16:03 UTC](https://meta.discourse.org/t/crawlers-very-high/151339/7 "2020-06-12T16:03:02Z")

</div>

This topic was automatically closed 30 days after the last reply. New replies are no longer allowed.
