# Oneboxing blocked by robot check

**URL:** https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594
**Category:** Support
**Created:** [December 17, 2020, 2:35pm UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594 "2020-12-17T14:35:25Z")
**Posts on this page:** 8
**Page:** 1

<div class="post-metadata">

### Author: ![breetai](https://avatars.discourse-cdn.com/v4/letter/b/a88e4f/32.png) [@breetai](https://meta.discourse.org/u/breetai)
#### Post date: [December 17, 2020, 2:35pm UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/1 "2020-12-17T14:35:25Z")

</div>

I am seeing this on a sites and it just started. When Discourse tries to pull the information from the site it is blocked. THis use to work in prior versions.  
I have included a link as an example

 ![image](https://global.discourse-cdn.com/meta/original/3X/a/a/aaa5f7c9141cc4cee1a00c0f58bc7f9b2b8bd058.png)

[Bloomberg - Are you a robot?](https://www.bloomberg.com/tosv2.html?vid=&uuid=dfb50010-4073-11eb-bde2-59e18427e7d0&url=L29waW5pb24vYXJ0aWNsZXMvMjAyMC0wMS0yOS9wZWVyLXJldmlldy1pcy1zY2llbmNlLXMtd2hlZWwtb2YtbWlzZm9ydHVuZQ==)

---

<div class="post-metadata">

### Author: ![pfaffman](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/pfaffman/32/120154_2.png) [@pfaffman](https://meta.discourse.org/u/pfaffman)
#### Post date: [December 17, 2020, 3:04pm UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/2 "2020-12-17T15:04:06Z")

</div>

It looks like rate limiting by Bloomberg. There likely isn’t much you can do other than infer what the limits are and contrive to stay under them.

---

<div class="post-metadata">

### Author: ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)
#### Post date: [December 19, 2020, 2:29am UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/3 "2020-12-19T02:29:58Z")

</div>

What exactly are you trying to onebox here? The URL is quite strange.

---

<div class="post-metadata">

### Author: ![breetai](https://avatars.discourse-cdn.com/v4/letter/b/a88e4f/32.png) [@breetai](https://meta.discourse.org/u/breetai)
#### Post date: [December 19, 2020, 2:42am UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/4 "2020-12-19T02:42:44Z")

</div>

Bloomberg news article. If you click on the link. Thahs the article.

---

<div class="post-metadata">

### Author: ![merefield](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/merefield/32/176214_2.png) [@merefield](https://meta.discourse.org/u/merefield)
#### Post date: [December 19, 2020, 7:39am UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/5 "2020-12-19T07:39:47Z")

</div>

Try ["Onebox Assistant", crawl for those previews reliably!](https://meta.discourse.org/t/onebox-assistant-crawl-for-those-previews-reliably/107405)

Works with Bloomberg links iirc.

What’s the original link? The one you’ve pasted above is not an article, it’s a destination you’ve been redirected to I suspect.

---

<div class="post-metadata">

### Author: ![breetai](https://avatars.discourse-cdn.com/v4/letter/b/a88e4f/32.png) [@breetai](https://meta.discourse.org/u/breetai)
#### Post date: [December 19, 2020, 7:34pm UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/6 "2020-12-19T19:34:20Z")

</div>

`https://www.bloomberg.com/opinion/articles/2020-01-29/peer-review-is-science-s-wheel-of-misfortune`

This is the link.

---

<div class="post-metadata">

### Author: ![codinghorror](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/codinghorror/32/110067_2.png) [@codinghorror](https://meta.discourse.org/u/codinghorror)
#### Post date: [December 19, 2020, 8:11pm UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/7 "2020-12-19T20:11:41Z")

</div>

I see, here’s the link

`http://www.bloomberg.com/opinion/articles/2020-01-29/peer-review-is-science-s-wheel-of-misfortune`

[http://www.bloomberg.com/opinion/articles/2020-01-29/peer-review-is-science-s-wheel-of-misfortune](http://www.bloomberg.com/opinion/articles/2020-01-29/peer-review-is-science-s-wheel-of-misfortune)

Looks like they have pretty aggressive anti-scraping in place, since all we’re doing is checking for metadata headers..

Also yet another example of where we shouldn’t be oneboxing at all because **we have neither image nor description** cc @techAPJ @sam .. we really gotta backport that change to stable once it goes in next week.

---

<div class="post-metadata">

### Author: ![JimPas](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/jimpas/32/148179_2.png) [@JimPas](https://meta.discourse.org/u/JimPas)
#### Post date: [December 20, 2020, 5:40am UTC](https://meta.discourse.org/t/oneboxing-blocked-by-robot-check/173594/8 "2020-12-20T05:40:22Z")

</div>

I just tried the link ending up to the html extension (minus all the trailing characters) just using Firefox, _not_ Discourse Onebox. The error extended error message is below the line. The 1st link (which has the error msg below) is enclosed in \<\> here. The second link is without being enclosed in \<\> and gives the title of the URL as shown.  
[https://www.bloomberg.com/tosv2.html](https://www.bloomberg.com/tosv2.html)  
[Bloomberg - Are you a robot?](https://www.bloomberg.com/tosv2.html)

* * *

## We’ve detected unusual activity from your computer network

To continue, please click the box below to let us know you’re not a robot.

### Why did this happen?

Please make sure your browser supports JavaScript and cookies and that you are not blocking them from loading. For more information you can review our [Terms of Service](https://www.bloomberg.com/notices/tos) and [Cookie Policy](https://www.bloomberg.com/notices/tos).

### Need Help?

For inquiries related to this message please [contact our support team](https://www.bloomberg.com/feedback) and provide the reference ID below.

Block reference ID: 13215fd0-4285-11eb-8faf-b7e9262e99b2
