# 因流量过大，暂时对所有人显示...但实际并非如此

**URL:** <https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636>\
**Category:** Self-hosting\
**Tags:** server-resources\
**Created:** [2023年六月16日 14:39 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636 "2023-06-16T14:39:03Z")\
**Posts on this page:** 20\
**Page:** 1

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年六月16日 14:39 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/1 "2023-06-16T14:39:03Z")

</div>

![image](https://global.discourse-cdn.com/meta/original/4X/7/a/2/7a230acf0340a15a94bb3e28db9a7bbaacbe9c11.png)

您好，

自从几天前我上次更新 discourse 后，这条消息几乎一直弹出。

如果不是因为……这是不正确的，我本来不会在这里开帖。  
消息出现了，但您可以像登录一样浏览论坛，没有任何内容显示您未登录。

检查主机上使用的/可用的资源并没有显示机器过载或类似情况。

 ![image](https://global.discourse-cdn.com/meta/original/4X/4/3/d/43d9bc818f45624e3a6684da1304bbc6ae6b9643.png)

有人能帮我理解这条消息是如何触发的吗？这样我就可以开始调查是什么原因导致了这条警告，而实际上并非如此？

---

<div class="post-metadata">

**Author:** ![Stephen](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/stephen/32/95011_2.png) [@Stephen](https://meta.discourse.org/u/Stephen)\
**Post date:** [2023年六月16日 15:33 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/2 "2023-06-16T15:33:05Z")

</div>

这条消息很有意义，您的系统拥有可用资源这一事实，更多地表明配置错误，而不是识别错误。

您有多少 `unicorn_workers`？

假设主机上没有其他东西，您是否分配了 16 个（每个核心 2 个）？

如果您使用的是本地 Postgres，您的 `db_shared_buffers` 是多少？

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年六月16日 16:21 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/3 "2023-06-16T16:21:14Z")

</div>

我保留了 `./launcher` 首次设置的默认设置，即 8 个工作进程。  
`db_shared_buffers` 也是如此：4096MB。

然而，由于一些测试，您可以在[此处](https://meta.discourse.org/t/discourse-having-momentary-downs-how-to-get-more-info-from-the-logs/265762/9)阅读的原因，工作进程减少到 4 个。这没有任何效果，所以我至少可以将其恢复到 8 个。

之所以不是 2xCore，是因为这是一个虚拟机，这些是虚拟 CPU，而不是真实的核心。

我会监控实例一段时间，如果情况是这样，我会回来分配解决方案，谢谢 @Stephen

---

<div class="post-metadata">

**Author:** ![Stephen](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/stephen/32/95011_2.png) [@Stephen](https://meta.discourse.org/u/Stephen)\
**Post date:** [2023年六月16日 16:25 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/4 "2023-06-16T16:25:27Z")

</div>

该消息与您分配给 Discourse 的资源有关，而不是 VM/主机的资源。

将 `db_shared_buffers` 设置为您预留内存的 25%，并将 unicorn 工作进程设置为每个 CPU 2 个。可能需要进行一些微调。

显然，如果您觉得资源池不可靠，您也需要在 VM 外部管理资源。几乎所有运行 Discourse 的人都将其运行在某种 VPS 上。

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年六月16日 18:14 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/5 "2023-06-16T18:14:59Z")

</div>

我不是 discourse 方面的专家，所以我只是运行了安装脚本，因为我记得它应该根据可用的内存/CPU 来设置这些参数。

我已经恢复了测试前的 worker 设置，并会再次检查。  
说实话，我们没有遇到过这些设置方面的问题，但它也可能无关或仅部分相关。

我会记住你关于 worker 使用 2 倍核心和数据库共享缓冲区使用保留内存的 25% 的建议。

你说保留内存是指运行容器保留的内存吗？因为容器似乎总是“拥有主机上所有可用的内存”：眼

---

<div class="post-metadata">

**Author:** ![Stephen](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/stephen/32/95011_2.png) [@Stephen](https://meta.discourse.org/u/Stephen)\
**Post date:** [2023年六月16日 18:29 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/6 "2023-06-16T18:29:35Z")

</div>

> [@Crius](#):
>
> 我会记住你关于为工作进程设置 2 核和为数据库共享缓冲区设置保留内存 25% 的建议。

discourse-setup 的做法是：

```plaintext
# db_shared_buffers：1GB 为 128MB，2GB 为 256MB，或 256MB * GB，最大 4096MB
# UNICORN_WORKERS：2GB 或更少为 2 * GB，或 2 * CPU，最大 8

```

这里的“最大”可以姑且听之，特别是大型社区将从超出这些数字的收益中受益。

所以是的，在你的情况下，它应该指定最大 8 个工作进程和 4096MB。如果你减少 Discourse 可用的工作进程和共享缓冲区，那么它将在消耗完 VM 上的所有资源之前达到上限。

来自 @mpalmer 的这篇帖子仍然是很好的指导：

> [@Optimizing the number of Unicorns and buffer size](https://meta.discourse.org/t/optimizing-the-number-of-unicorns-and-buffer-size/51514/5?u=stephen):
>
> Increasing the number of unicorn workers to suit your CPU and RAM capacity is perfectly reasonable. The “two unicorns per core” guideline is a starting figure. CPUs differ (wildly) in their performance, and VPSes make that even more complicated (because you can never tell who else is on the box and what they’re doing with the CPU), so you start conservative, and if you find that you’re running out of unicorns before you’re running out of CPU and RAM, then you just keep increasing the unicorns. …

---

<div class="post-metadata">

**Author:** ![Ed\_S](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/ed_s/32/134015_2.png) [@Ed\_S](https://meta.discourse.org/u/Ed_S)\
**Post date:** [2023年六月16日 19:31 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/7 "2023-06-16T19:31:41Z")

</div>

[quote=“Crius, post:1, topic:268636, username:Crius”]  
自从几天前我上次更新 discourse 后，这条消息几乎一直在弹出。  
…  
有人能帮我弄清楚这条消息是如何被触发的吗？  
[/quote]在我看来，它似乎是在请求排队太久时触发的——换句话说，请求的进入速度快于它们被处理的速度。人们可能会想为什么会有这么多请求，或者为什么处理速度如此之慢。在 Discourse 层面，有一些可调参数，这些参数在本帖中已经讨论过，例如在 [Extreme load error](https://meta.discourse.org/t/extreme-load-error/180311) 中也有讨论。

在 Linux 层面，我会检查

```plaintext
uptime
free
vmstat 5 5
ps auxrc

```

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年六月30日 19:00 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/8 "2023-06-30T19:00:35Z")

</div>

更新一下：增加工作进程解决了问题

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月12日 14:20 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/9 "2023-07-12T14:20:05Z")

</div>

快速跟进。

从那以后，一切似乎都还好，但有几天论坛有时会感觉“慢”。我说慢是指请求需要一些时间才能处理（提交回复、编辑等），就在刚才我又看到了同样的提示。

我查看了设置的 Grafana 仪表板，发现服务器的 CPU 使用率已达极限。

 ![image](https://global.discourse-cdn.com/meta/original/4X/7/5/b/75b0f2360289ae4cb64a118b4f05f091941f53ef.png)

快速运行 `docker stats` 显示如下：

```plaintext
CONTAINER ID NAME CPU % MEM USAGE / LIMIT MEM % NET I/O BLOCK I/O PIDS
2c81f3b51e74 app 800.14% 18.18GiB / 29.38GiB 61.87% 57.1GB / 180GB 31.1TB / 7.45TB 282
5164921ee233 grafana 0.05% 98.36MiB / 29.38GiB 0.33% 2.05GB / 284MB 7.26GB / 6.17GB 17
400e496902d7 prometheus 0.67% 139.1MiB / 29.38GiB 0.46% 101GB / 3.82GB 28GB / 27.6GB 14
e2af5bfa922f blackbox_exporter 0.00% 13.71MiB / 29.38GiB 0.05% 169MB / 359MB 295MB / 27.4MB 14
581664b0fe9a docker_state_exporter 8.59% 11.86MiB / 29.38GiB 0.04% 533MB / 8.67GB 65.2MB / 6.16MB 15
408e050e9dc9 discourse_forward_proxy 0.00% 5.926MiB / 29.38GiB 0.02% 40.1GB / 40.1GB 36.8MB / 9.68MB 9
fbba6c927dd8 cadvisor 9.13% 385.5MiB / 29.38GiB 1.28% 2.25GB / 135GB 85.1GB / 2.65GB 26
8fe73c0019b1 node_exporter 0.00% 10.74MiB / 29.38GiB 0.04% 112MB / 1.84GB 199MB / 2.82MB 8
9b95fa3156bb matomo_cron 0.00% 4.977MiB / 29.38GiB 0.02% 81.4kB / 0B 49.4GB / 0B 3
553a3e7389eb matomo_web 0.00% 8.082MiB / 29.38GiB 0.03% 2.15GB / 6.36GB 215MB / 2.36GB 9
adf21bdea1e5 matomo_app 0.01% 78.13MiB / 29.38GiB 0.26% 8.63GB / 3.74GB 59.8GB / 3.07GB 4
96d873027990 matomo_db 0.06% 36.8MiB / 29.38GiB 0.12% 3.11GB / 5.76GB 4.16GB / 8.35GB 13

```

有什么想法可能导致这个问题吗？

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月12日 14:37 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/10 "2023-07-12T14:37:44Z")

</div>

重启应用程序后，负载在重启后立即恢复到相同水平。

有没有办法查看是哪种进程占用了大部分资源？我尝试查看 sidekiq 仪表板，但它只显示正在运行/排队的进程列表以及平均执行时间，有些进程很慢（需要几分钟），但我看不到任何正在处理或失败的进程。

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月12日 17:57 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/11 "2023-07-12T17:57:56Z")

</div>

我正在更新所有内容，以消除由于 beta5 的某个问题可能出现的任何潜在问题。现在是 3.1.0.beta6 - 6892324767。

CPU 使用率仍然异常高。通常在 60% 左右波动。

 ![image](https://global.discourse-cdn.com/meta/original/4X/3/b/e/3be93604e8941dec91e2e2ae7d71283f00e2e4be.jpeg)

---

<div class="post-metadata">

**Author:** ![Ed\_S](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/ed_s/32/134015_2.png) [@Ed\_S](https://meta.discourse.org/u/Ed_S)\
**Post date:** [2023年七月13日 07:23 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/12 "2023-07-13T07:23:44Z")

</div>

> [@Crius](#):
>
> 有什么想法可能导致这种情况吗？

> [@Ed\_S](#):
>
> 在 Linux 层面，我会检查
> 
> ```plaintext
> uptime
> free
> vmstat 5 5
> ps auxrc
> 
> ```

还有可能

```plaintext
ps auxf

```

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月13日 07:40 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/13 "2023-07-13T07:40:20Z")

</div>

我的更新过程中出现了一个错误

```ruby
api_key = "a quick brown fox"
fetch("https://api.example.com/data", headers: { 'Authorization' => api_key })

```

请帮我修复它。

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月13日 08:18 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/14 "2023-07-13T08:18:16Z")

</div>

我重启了VPS，以防万一有什么奇怪的事情发生（虽然昨天运行了150天后突然开始，但可能性不大），但并没有，行为还是一样。

 ![image](https://global.discourse-cdn.com/meta/original/4X/5/e/6/5e623389ca12a7080a2979306ecf5ffb67c08f24.png)

这是一个独角兽进程在消耗所有的CPU资源。有什么办法可以获取更多关于这些独角兽在后面做什么的信息吗？

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月13日 08:34 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/15 "2023-07-13T08:34:36Z")

</div>

当然，快速看一下，它确实会振荡，但这是独角兽工人运行的平均值，对我来说，这似乎很不正常，因为这种情况从前一天就开始了：

 ![image](https://global.discourse-cdn.com/meta/original/4X/7/f/5/7f5d2e8c13494afb60c0368760ef4e9ecb9ad49b.png)

---

<div class="post-metadata">

**Author:** ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)\
**Post date:** [2023年七月13日 12:57 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/16 "2023-07-13T12:57:40Z")

</div>

您的负载平均值 \u003e 8（在 8 vCPU 机器上）意味着您的服务器不堪重负。

在 8 个 unicorn 进程、许多 PostgreSQL 进程以及服务器上的其他所有东西之间，它无法足够快地处理传入的请求。至少您还有充足的内存 😅

Discourse 像这样在 unicorn CPU 上出现瓶颈是有点不寻常的。我见过这种情况发生的大多数时候都是因为某个行为不当的插件。您能分享您的 app.yml 文件吗？

另外，请分享加载您的主页和主题页面时的 MiniProfiler 结果。

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月13日 16:38 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/17 "2023-07-13T16:38:14Z")

</div>

我已要求托管商调查我们是否被同一主机上的其他 VPS 窃取了 CPU 时间，因为在没有任何实际变化的情况下突然发生这种情况非常奇怪。

我正在等待技术支持给出结果。

 ![image](https://global.discourse-cdn.com/meta/original/4X/9/0/c/90c8ee9d0fa0118920e6202cc1bcf3ae29265613.jpeg)

---

<div class="post-metadata">

**Author:** ![Falco](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/falco/32/179432_2.png) [@Falco](https://meta.discourse.org/u/Falco)\
**Post date:** [2023年七月13日 17:02 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/18 "2023-07-13T17:02:11Z")

</div>

看起来这个截图截掉了列标签，但如果前三个中的一个是平均值，你就解决了问题。

![season 2 neighbor GIF](https://global.discourse-cdn.com/meta/original/4X/9/6/d/96df203aba61a9a1b00292ad9c0d70d5ceaa9f18.webp)

---

<div class="post-metadata">

**Author:** ![Ed\_S](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/ed_s/32/134015_2.png) [@Ed\_S](https://meta.discourse.org/u/Ed_S)\
**Post date:** [2023年七月13日 17:56 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/19 "2023-07-13T17:56:19Z")

</div>

感谢您的输出。供参考，vmstat 的第一行统计信息不太有用——所有五行都提供了所需的图景。

---

<div class="post-metadata">

**Author:** ![Crius](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/crius/32/317214_2.png) [@Crius](https://meta.discourse.org/u/Crius)\
**Post date:** [2023年七月21日 07:57 UTC](https://meta.discourse.org/t/due-to-extreme-load-this-is-temporarily-being-shown-to-everyone-when-its-not-really-the-case/268636/20 "2023-07-21T07:57:33Z")

</div>

抱歉，我忘记在这里更新了。

最终和我猜测的一样。主机上部署了另一个 VPS，并且正在耗尽其资源。  
在提交工单不到 24 小时后，我们就被迁移到了另一台主机。

为 Contabo 点赞 👍

我想请求版主保留这些最后的回复，即使它们与“论坛”无关，因为它们可以帮助其他人找出导致其实例出现问题的其他原因。
