# 为什么 robots.txt 中有很多 Disallow 规则？

**URL:** <https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251>\
**Category:** Support\
**Created:** [2018年三月30日 19:23 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251 "2018-03-30T19:23:27Z")\
**Posts on this page:** 15\
**Page:** 2

<div class="post-metadata">

**Author:** ![sam](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/sam/32/102149_2.png) [@sam](https://meta.discourse.org/u/sam)\
**Post date:** [2018年三月30日 21:45 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/22 "2018-03-30T21:45:10Z")

</div>

Well bing has a bug and is not respecting robots txt, why not raise this with Microsoft

Note, I am sure some real weird things are up with bing, looking at crawler stats here it’s going ballistic on meta

---

<div class="post-metadata">

**Author:** ![notriddle](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/notriddle/32/133055_2.png) [@notriddle](https://meta.discourse.org/u/notriddle)\
**Post date:** [2018年三月30日 21:47 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/24 "2018-03-30T21:47:31Z")

</div>

Actually, it is respecting robots.txt. It is not crawling the pages that are denied. It is including the page in its index, but the robots.txt spec allows you to do that just fine.

---

<div class="post-metadata">

**Author:** ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)\
**Post date:** [2018年三月30日 21:48 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/26 "2018-03-30T21:48:45Z")

</div>

> [@sam](#):
>
> some real weird things are up with bing, looking at crawler stats here it’s going ballistic on meta

I refer you to my previous statements on the matter:

> [@codinghorror](#):
>
> Bing is terrible

We saw this at Stack Overflow as well. A lot of sound and fury from bing crawlers resulting in virtually zero real traffic. Bing completely sucks, and has for a decade.

---

<div class="post-metadata">

**Author:** ![Gulshan\_Kumar](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/gulshan_kumar/32/119562_2.png) [@Gulshan\_Kumar](https://meta.discourse.org/u/Gulshan_Kumar)\
**Post date:** [2018年三月30日 21:50 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/27 "2018-03-30T21:50:09Z")

</div>

> [@sam](#):
>
> Well bing has a bug and is not respecting robots txt, why not raise this with Microsoft

The disallow rule at robots.txt file is to prevent from crawling, never from indexing.  
Raising the issue with MS? Seriously, they don’t care about webmasters.

---

<div class="post-metadata">

**Author:** ![sam](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/sam/32/102149_2.png) [@sam](https://meta.discourse.org/u/sam)\
**Post date:** [2018年三月30日 21:52 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/28 "2018-03-30T21:52:09Z")

</div>

I am against removing this from robots, but somehow amending discourse to carry the list and auto generate don’t index meta tags as well on those pages is something I am open to, especially if this reduces traffic from bing

---

<div class="post-metadata">

**Author:** ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)\
**Post date:** [2018年三月30日 21:53 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/29 "2018-03-30T21:53:06Z")

</div>

> [@sam](#):
>
> especially if this reduces traffic from bing

Worth a try, at least _that_ is a net benefit to every Discourse instance.

---

<div class="post-metadata">

**Author:** ![notriddle](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/notriddle/32/133055_2.png) [@notriddle](https://meta.discourse.org/u/notriddle)\
**Post date:** [2018年三月30日 21:56 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/31 "2018-03-30T21:56:44Z")

</div>

Personally, @Gulshan_Kumar, I recommend:

1. Work on getting people to link to your forum’s front page from Twitter and stuff. Unindexed pages should not be ranking higher than your front page, even on Bing.

2. Don’t worry about Bing returning your user profile when searching for your login name. That’s correct behavior.

And as for @codinghorror: don’t act like Stack Overflow is representative of the internet at large. Bing, being the default search engine in Internet Explorer, gets most of its traffic from a very different demographic than SO or even Discourse Meta targets.

---

<div class="post-metadata">

**Author:** ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)\
**Post date:** [2018年三月30日 21:58 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/32 "2018-03-30T21:58:32Z")

</div>

> **[Bing search market share by country 2017| Statista](https://www.statista.com/statistics/220538/bing-search-market-share-country/)**
>
> This statistic shows the worldwide search market share of Bing as of August 2017 in leading online markets.

Worldwide, 9%. My main beef is that Bing is objectively _very bad_. Both at providing relevant results, and the crawler behavior.

I mean for a criteria of “type words in a search box and have _something_ come back”, it works..

---

<div class="post-metadata">

**Author:** ![Gulshan\_Kumar](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/gulshan_kumar/32/119562_2.png) [@Gulshan\_Kumar](https://meta.discourse.org/u/Gulshan_Kumar)\
**Post date:** [2018年三月30日 22:06 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/33 "2018-03-30T22:06:47Z")

</div>

About Bing, I think every bit helps. If a website is quality enough, at least for some keywords it will rank.  
My purpose to keep the `noindex` is to avoid indexing unnecessary pages. The way Bing/Yahoo randomly show stuff, I feel helpless.

Thanks for participating in this discussion. I greatly appreciate your valuable input.

> [@notriddle](#):
>
> Also, user profiles contain excerpts of the posts, but we don’t want the search engines linking there. They should be linking to the actual posts.

Sorry, I forget to mention. How if we just remove the user-profile link for non-logged in users (anonymous/bot)? This will solve all problem.

---

<div class="post-metadata">

**Author:** ![notriddle](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/notriddle/32/133055_2.png) [@notriddle](https://meta.discourse.org/u/notriddle)\
**Post date:** [2018年三月30日 22:26 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/36 "2018-03-30T22:26:01Z")

</div>

Because they’re not supposed to be secret.

---

<div class="post-metadata">

**Author:** ![riking](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/riking/32/170938_2.png) [@riking](https://meta.discourse.org/u/riking)\
**Post date:** [2018年三月30日 22:32 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/37 "2018-03-30T22:32:08Z")

</div>

Also note that the crawler view user profile… doesn’t actually contain anything other than the bio. Certainly not a list of links to posts.

Now that the user page has been redefined a few times, presenting these links to crawlers might actually be a good thing - surfacing quality content is what those sections are for, and we might as well help the search engine out with them.

@sam’s profile Overview

 ![image](https://global.discourse-cdn.com/meta/original/3X/9/4/9498363d234c6ebe5cd75a6d74b3d71fc2109eef.png)

versus noscript version of me:

```plaintext

      <div id="main-outlet" class="wrap">
        <!-- preload-content: -->
         <h2>riking</h2>

<p><p>Discourse is pretty great</p>
<p><a href="https://github.com/riking" class="onebox" target="_blank">https://github.com/riking</a><br>
<a href="https://twitter.com/riking27" class="onebox" target="_blank">https://twitter.com/riking27</a></p></p>

        <!-- :preload-content -->
        <footer>
          <nav itemscope itemtype='http://schema.org/SiteNavigationElement'>
            <a href='/'>Home</a>
            <a href="/categories">Categories</a>
            <a href="/guidelines">FAQ/Guidelines</a>
            <a href="/tos">Terms of Service</a>
            <a href="/privacy">Privacy Policy</a>
          </nav>
        </footer>
      </div>

      <footer id='noscript-footer'>
        <p>Powered by <a href="https://www.discourse.org">Discourse</a>, best viewed with JavaScript enabled</p>
      </footer>

```

_wow_, that’s out of date…

---

<div class="post-metadata">

**Author:** ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)\
**Post date:** [2018年四月1日 07:45 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/39 "2018-04-01T07:45:47Z")

</div>

See also

> **[U.S. search engine market share queries handled 2025| Statista](https://www.statista.com/statistics/267161/market-share-of-search-engines-in-the-united-states/)**
>
> In February 2025, Microsoft Sites handled \*\*\*\* percent of all search queries in the United States.

Google has a massive lead on mobile as well, so that transition heavily favors them.

---

<div class="post-metadata">

**Author:** ![markersocial](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/markersocial/32/170136_2.png) [@markersocial](https://meta.discourse.org/u/markersocial)\
**Post date:** [2018年七月11日 22:43 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/40 "2018-07-11T22:43:14Z")

</div>

> [@notriddle](#):
>
> Because they’re infrequently read by humans, [black hat SEO](https://en.wikipedia.org/wiki/Search_engine_optimization) people like to spam forum user profiles. Discourse blocks user profiles in order to disincentivise it.

@notriddle - I think profiles are quite frequently read by humans, forum profiles will often have more views than any one of the user’s threads. If blocking user profiles from indexing is to disincentivize spammers from placing links, it means that we’re likely in agreement that links listed inside a profile will be valued more if the page is indexed, as opposed to not indexed.

This would mean that our internal linking will suffer, due to being on noindex profile pages. More strong internal linking is better for SEO and helps search engines crawl content better. Especially when no sitemap is being used. Simply adding nofollow to external spam/non-whitelisted profile links should be sufficient to disincentivize spammers (if it isn’t done already), many probably don’t even check the robots.txt to see if the profiles get indexed.

Here is a thread incl. video discussing how Google will not follow links on noindex pages: [Google Will Eventually Stop Following Links on Noindex Pages - Google Search and SEO forum at WebmasterWorld - WebmasterWorld](https://www.webmasterworld.com/google/4881752.htm)

> [@codinghorror](#):
>
> That’s not the only reason; the user profiles are all duplicate data, in that your posts and topics are visible on… the topics themselves, which _are_ the focus.

@codinghorror - As for it being duplicate content, I think profiles are less duplicate content than a thread being inside parent and child categories (incl. tags) simultaneously. The links are duplicates, but on different pages/urls, with different purposes. Profiles can include unique content also like bio, the floating card that also displays the bio displays on desktop only, to view it on mobile you can only go to their full profile. Google has switched to mobile first indexing also, meaning the mobile versions of our sites have become the primary versions of our sites: [How Does Mobile-First Indexing Work, and How Does It Impact SEO? - Moz](https://moz.com/blog/mobile-first-indexing-seo)

The question is, why wouldn’t we want to allow search engines to have better crawling on our sites? Reddit allows indexing of user profiles and is basically the largest forum in the world. Youtube allows indexing of channels, Twiiter, FB, Google Plus etc. allow indexing of profiles as well. The only main examples I’ve seen of using noindex for profiles is on old forum software.

I definitely think the noindex for user profile pages in the robots.txt should not be default.

---

<div class="post-metadata">

**Author:** ![sam](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/sam/32/102149_2.png) [@sam](https://meta.discourse.org/u/sam)\
**Post date:** [2020年十二月22日 02:50 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/41 "2020-12-22T02:50:32Z")

</div>

只是重新激活这个讨论。

1. 现在，如果您愿意，可以随意编辑 robots.txt 文件。

2. 对于不应被索引的页面，我们始终提供 `x-robots-tag` noindex 标签。

3. 事实证明，如果不在 robots.txt 中提供严格指导，某些爬虫会对网站“大肆扫荡”，并非所有爬虫都像 Google 那样守规矩。如今我们的 robots.txt 文件非常基础，但这付出了代价。（我们期望所有爬虫都能像 Google 一样守规矩，但要成为 Google 需要付出巨大努力。）

我认为我们至少应该默认恢复对所有非 Googlebot 的“非常严格”的 robots.txt 设置。

---

<div class="post-metadata">

**Author:** ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)\
**Post date:** [2020年十二月22日 04:28 UTC](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251/42 "2020-12-22T04:28:14Z")

</div>

> [@sam](#):
>
> 我认为我们至少应该默认恢复“非常严格”的 robots.txt，对所有非 Googlebot 的爬虫生效。

当然，请随意这样设置……Google 某种程度上把我们逼到了这个境地。

[上一頁](https://meta.discourse.org/t/why-there-are-lots-of-disallow-rule-in-robots-txt/84251.md?page=1)
