# 改进如何检测并标记流量数据中可能的爬虫

**URL:** https://meta.discourse.org/t/improving-how-we-detect-and-flag-likely-crawlers-in-your-traffic-data/409316
**Category:** Announcements
**Tags:** dashboard
**Created:** [2026年八月20日 01:16 UTC](https://meta.discourse.org/t/improving-how-we-detect-and-flag-likely-crawlers-in-your-traffic-data/409316 "2026-08-20T01:16:10Z")
**Posts on this page:** 1
**Showing post:** 10

<div class="post-metadata">

### Author: ![kris.kotlarek](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/kris.kotlarek/32/176919_2.png) [@kris.kotlarek](https://meta.discourse.org/u/kris.kotlarek)
#### Post date: [2026年八月21日 03:15 UTC](https://meta.discourse.org/t/improving-how-we-detect-and-flag-likely-crawlers-in-your-traffic-data/409316/10 "2026-08-21T03:15:59Z")

</div>

> [@nickdb](#):
>
> 如果能获取一份它已识别的用户代理（user agent）列表，或者它认为可能是机器人/爬虫的列表，那就太好了。

难点在于，所谓的“疑似爬虫”往往使用的是合法的用户代理。有时浏览器版本看起来比较旧，但这不足以作为判定其为疑似机器人的强信号。因此，我们收集了大量指标，并尝试将它们结合起来以评估可能性。

当爬虫诚实地通过类似 Bingbot 或 Googlebot 这样的用户代理表明身份时，它会被直接归类到“爬虫”类别中。

---

_[View the full topic](https://meta.discourse.org/t/improving-how-we-detect-and-flag-likely-crawlers-in-your-traffic-data/409316)._
