# Sidekiq hangs (on BotInput job?)

**URL:** https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661
**Category:** Support
**Tags:** discobot
**Created:** [February 17, 2025, 8:21am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661 "2025-02-17T08:21:30Z")
**Posts on this page:** 17
**Page:** 1

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [February 17, 2025, 8:21am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/1 "2025-02-17T08:21:30Z")

</div>

In the past week we have seen three Sidekiq instances on different forums being stuck. There was nothing special going on, it was just that Sidekiq was not processing any work and showing 5 of 5 jobs being processed.

One interesting thing they all had in common was that there was one critical `BotInput` job among the jobs. Now this is quite a common job, but it still stands out.

After restarting Sidekiq everything works normal again. Manually queuing a job with the same parameters does not cause it to hang again. There is nothing special with the specific post it was called for.

Does anyone have any idea how we could track down what is going on here?

---

<div class="post-metadata">

### Author: ![Shauny](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/shauny/32/362012_2.png) [@Shauny](https://meta.discourse.org/u/Shauny)
#### Post date: [February 17, 2025, 9:34pm UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/2 "2025-02-17T21:34:19Z")

</div>

We have also been having hangs like this, and our host can’t figure out what is causing it.

---

<div class="post-metadata">

### Author: ![tgxworld](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tgxworld/32/106117_2.png) [@tgxworld](https://meta.discourse.org/u/tgxworld)
#### Post date: [February 18, 2025, 1:30am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/3 "2025-02-18T01:30:57Z")

</div>

> [@RGJ](#):
>
> it was just that Sidekiq was not processing any work and showing 5 of 5 jobs being processed.

Do you have a screenshot of what you are seeing in the dashboard?

If you can, please try sending the Sidekiq process the `TTIN` signal and provide the backtrace here.

> **[Signals](https://github.com/sidekiq/sidekiq/wiki/Signals#ttin)**
>
> Simple, efficient background processing for Ruby. Contribute to sidekiq/sidekiq development by creating an account on GitHub.

---

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [February 20, 2025, 8:05am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/4 "2025-02-20T08:05:31Z")

</div>

Sorry, took a while before this happened again.

 ![sidekiq-20250220](https://global.discourse-cdn.com/meta/original/4X/7/6/1/761e0205df444d6f370e12b8a2963b9d95f90f4a.png)

[sidekiq-clean.txt](https://meta.discourse.org/uploads/short-url/aTiGa9g7rgh4apfotfDKPPL8IoU.txt) (35.8 KB)

Summary of the logs

```plaintext
[default] Thread TID-1ow77 
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/sidekiq-6.5.12/lib/sidekiq/cli.rb:199:in `backtrace'
--
[default] Thread TID-1o1jr 
[default] /var/www/discourse/lib/demon/base.rb:234:in `sleep'
--
[default] Thread TID-1o1j7 
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/redis-4.8.1/lib/redis/connection/ruby.rb:57:in `wait_readable'
--
[default] Thread TID-1o1j3 
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/message_bus-4.3.8/lib/message_bus/timer_thread.rb:130:in `sleep'
--
[default] Thread TID-1o1ij AR Pool Reaper
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/activerecord-7.2.2.1/lib/active_record/connection_adapters/abstract/connection_pool/reaper.rb:49:in `sleep'
--
[default] Thread TID-1o1hj 
[default] <internal:thread_sync>:18:in `pop'
--
[default] Thread TID-1o1gz AR Pool Reaper
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/activerecord-7.2.2.1/lib/active_record/connection_adapters/abstract/connection_pool/reaper.rb:49:in `sleep'
--
[default] Thread TID-1o1gv 
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/mini_scheduler-0.18.0/lib/mini_scheduler/manager.rb:18:in `sleep'
--
[default] Thread TID-1o1gb 
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/mini_scheduler-0.18.0/lib/mini_scheduler/manager.rb:32:in `sleep'
--
[default] Thread TID-1otmb 
[default] /var/www/discourse/app/models/top_topic.rb:8:in `refresh_daily!'
--
[default] Thread TID-1otkn 
[default] /var/www/discourse/app/models/top_topic.rb:8:in `refresh_daily!'
--
[default] Thread TID-1otjz 
[default] /var/www/discourse/app/models/top_topic.rb:8:in `refresh_daily!'
--
[default] Thread TID-1otif 
[default] /var/www/discourse/app/models/top_topic.rb:8:in `refresh_daily!'
--
[default] Thread TID-1othr 
[default] /var/www/discourse/app/models/top_topic.rb:8:in `refresh_daily!'
--
[default] Thread TID-1o1fb 
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/mini_scheduler-0.18.0/lib/mini_scheduler.rb:80:in `sleep'
--
[default] Thread TID-1o1er 
[default] /var/www/discourse/lib/mini_scheduler_long_running_job_logger.rb:87:in `sleep'
--
[default] Thread TID-1o1en heartbeat
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/sidekiq-6.5.12/lib/sidekiq/launcher.rb:76:in `sleep'
--
[default] Thread TID-1o1e3 scheduler
[default] /var/www/discourse/vendor/bundle/ruby/3.3.0/gems/connection_pool-2.5.0/lib/connection_pool/timed_stack.rb:79:in `sleep'
--
[default] Thread TID-1ot4n processor
[default] /var/www/discourse/app/models/email_log.rb:58:in `unique_email_per_post'
--
[default] Thread TID-1ot67 processor
[default] /var/www/discourse/app/models/email_log.rb:58:in `unique_email_per_post'
--
[default] Thread TID-1ot8j processor
[default] /var/www/discourse/app/models/email_log.rb:58:in `unique_email_per_post'
--
[default] Thread TID-1ot5n processor
[default] /usr/local/lib/ruby/3.3.0/bundled_gems.rb:74:in `require'
--
[default] Thread TID-1ot6b processor
[default] /var/www/discourse/lib/distributed_mutex.rb:5:in `<main>'
--
[default] Thread TID-1o0kn final-destination_resolver_thread
[default] <internal:thread_sync>:18:in `pop'
--
[default] Thread TID-1o0k3 Timeout stdlib thread
[default] <internal:thread_sync>:18:in `pop'

```

---

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [March 3, 2025, 7:57am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/5 "2025-03-03T07:57:47Z")

</div>

@tgxworld did you have a chance to look at the backtrace?

---

<div class="post-metadata">

### Author: ![Isambard](https://avatars.discourse-cdn.com/v4/letter/i/858c86/32.png) [@Isambard](https://meta.discourse.org/u/Isambard)
#### Post date: [March 3, 2025, 9:08am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/7 "2025-03-03T09:08:53Z")

</div>

I have been having Sidekiq issues since a forum upgrade a month ago. What command do you use to restart Sidekiq? Just a `sv restart sidekiq`?

---

<div class="post-metadata">

### Author: ![tgxworld](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tgxworld/32/106117_2.png) [@tgxworld](https://meta.discourse.org/u/tgxworld)
#### Post date: [March 5, 2025, 1:17am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/8 "2025-03-05T01:17:54Z")

</div>

Sorry I have not had a chance to take a look yet. Will try and get to it sometime this week.

---

<div class="post-metadata">

### Author: ![markschmucker](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/markschmucker/32/141599_2.png) [@markschmucker](https://meta.discourse.org/u/markschmucker)
#### Post date: [March 6, 2025, 7:46pm UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/9 "2025-03-06T19:46:52Z")

</div>

I’m seeing this in the past few days. Eventually all jobs stop running. Previously I rebooted, but is it safe to delete the critical queue? Is it a redis queue?

I’m up-to-date at 3.5.0.beta1-dev.

Just a wild guess, but sometimes when I’m chatting with the bot it stops responding so I refresh the page or give up. Maybe those cases leave a job hanging?

 ![image](https://global.discourse-cdn.com/meta/original/4X/8/a/0/8a0e4e4e1ac4b993483e0dd6d4b59b8756fe4a40.png)

 ![image](https://global.discourse-cdn.com/meta/original/4X/4/5/5/4555b122e86ccd5459497960f3fd746abce7112e.png)

---

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [March 6, 2025, 8:21pm UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/10 "2025-03-06T20:21:58Z")

</div>

> [@markschmucker](#):
>
> when I’m chatting with the bot it stops responding so I refresh the page or give up. Maybe those cases leave a job hanging?

These jobs are asynchronous so they wouldn’t even know that you did that.

It’s interesting to hear that you are having this on `Jobs::BotInput` as well. We’re seeing this issue on only a small subset of all our servers (a few percent) and it seems to be the instances that use the narrative bot quite heavily.

> [@markschmucker](#):
>
> is it safe to delete the critical queue?

No, you would lose all the other queued jobs as well.

The most easy and safe way is `sv reload unicorn` from within the container.

---

<div class="post-metadata">

### Author: ![markschmucker](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/markschmucker/32/141599_2.png) [@markschmucker](https://meta.discourse.org/u/markschmucker)
#### Post date: [March 7, 2025, 1:39am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/11 "2025-03-07T01:39:16Z")

</div>

> [@RGJ](#):
>
> We’re seeing this issue on only a small subset of all our servers (a few percent) and it seems to be the instances that use the narrative bot quite heavily.

That’s not the case with our forum. AI is only visible to staff and I’ve confirmed no staffers are using it.

I’ve disabled AI for now.

---

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [March 7, 2025, 8:05am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/12 "2025-03-07T08:05:36Z")

</div>

`BotInput` is a job from the Discourse Narrative Bot (aka Discobot), not the AI bot.

---

<div class="post-metadata">

### Author: ![markschmucker](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/markschmucker/32/141599_2.png) [@markschmucker](https://meta.discourse.org/u/markschmucker)
#### Post date: [March 7, 2025, 10:45am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/13 "2025-03-07T10:45:22Z")

</div>

Ah. I have been using the API heavily, as the username discobot.

---

<div class="post-metadata">

### Author: ![tgxworld](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tgxworld/32/106117_2.png) [@tgxworld](https://meta.discourse.org/u/tgxworld)
#### Post date: [March 10, 2025, 8:36am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/14 "2025-03-10T08:36:34Z")

</div>

I had a look at the backtraces and it all points to some problem with the following line:

> <https://github.com/discourse/discourse/blob/85e525a8d7545c8b76966a60d72c4547f22a314e/plugins/discourse-narrative-bot/lib/discourse_narrative_bot/new_user_narrative.rb>

Not exactly sure why that line would cause problems though but it is a line that is not necessary so I’ve dropped it in

[https://github.com/discourse/discourse/pull/31727](https://github.com/discourse/discourse/pull/31727)

@RGJ Do you happen to have `Rails.application.config.eager_load` set to disable for some reason? 🤔

---

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [March 10, 2025, 8:58am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/16 "2025-03-10T08:58:10Z")

</div>

Interesting find, thank you for looking into it.  
It’s hard to tell when such an intermittent problem goes away. I have removed that line on the three instances that hung the most often (one of them almost daily). I will check back in here either:

- when one of those instances hangs (we then know that this did not do the trick)
- on Friday if none of them hung (we can then start assuming it was the solution)

> [@tgxworld](#):
>
> Do you happen to have `Rails.application.config.eager_load` set to disable for some reason? 🤔

Nope, didn’t mess with that.

But someone did…

`Rails.autoloaders.main.do_not_eager_load(config.root.join("lib"))`

at [Blaming discourse/config/application.rb at main · discourse/discourse · GitHub](https://github.com/discourse/discourse/blame/main/config/application.rb#L110)

---

<div class="post-metadata">

### Author: ![tgxworld](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/tgxworld/32/106117_2.png) [@tgxworld](https://meta.discourse.org/u/tgxworld)
#### Post date: [March 14, 2025, 1:11am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/17 "2025-03-14T01:11:24Z")

</div>

> [@RGJ](#):
>
> Rails.autoloaders.main.do\_not\_eager\_load(config.root.join(“lib”))

@loic Do you recall why we do not eager load the `lib` directory even in production?

---

<div class="post-metadata">

### Author: ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)
#### Post date: [March 14, 2025, 8:30am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/18 "2025-03-14T08:30:54Z")

</div>

> [@RGJ](#):
>
> It’s hard to tell when such an intermittent problem goes away. I have removed that line on the three instances that hung the most often (one of them almost daily). I will check back in here either:
> 
> - when one of those instances hangs (we then know that this did not do the trick)
> - on Friday if none of them hung (we can then start assuming it was the solution)

While the issues have been occuring this week, they haven’t been happening on the three instances where we removed that `require` line, so I think we can safely assume that this is the culprit 🎉 . Thank you for spotting that @tgxworld , I would have _never_ found that.

Would you be able to backport that fix to stable?

---

<div class="post-metadata">

### Author: ![loic](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/loic/32/105621_2.png) [@loic](https://meta.discourse.org/u/loic)
#### Post date: [March 14, 2025, 8:50am UTC](https://meta.discourse.org/t/sidekiq-hangs-on-botinput-job/352661/19 "2025-03-14T08:50:41Z")

</div>

> [@tgxworld](#):
>
> @loic Do you recall why we do not eager load the `lib` directory even in production?

It’s related to what’s explained here (when we upgraded to Rails 7.1): [Upgrading Ruby on Rails — Ruby on Rails Guides](https://guides.rubyonrails.org/upgrading_ruby_on_rails.html#config-autoload-lib-and-config-autoload-lib-once)

I don’t remember the exact problem, but we actually kept the previous behavior, having to require things from `lib`.
