# 由于数据库重连失败，恢复操作失败

**URL:** <https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615>\
**Category:** Bug\
**Created:** [2026年九月29日 23:40 UTC](https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615 "2026-09-29T23:40:34Z")\
**Posts on this page:** 1\
**Page:** 1

<div class="post-metadata">

**Author:** ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)\
**Post date:** [2026年九月29日 23:40 UTC](https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615/1 "2026-09-29T23:40:34Z")

</div>

### 情况描述：

一个由两个容器组成的 Discourse 安装环境，Web 和数据分别部署在不同的主机上。

### 问题

从非常旧的 Discourse 版本（即需要运行许多耗时的迁移任务）执行一个 22 GB 的大型命令行恢复操作时，在数据库恢复完成后立即失败，耗时 45 分钟。

```plaintext
Reconnecting to the database...
EXCEPTION: PQconsumeInput() could not receive data from server: Connection timed out
SSL SYSCALL error: Connection timed out
/var/www/discourse/vendor/bundle/ruby/3.4.0/gems/rack-mini-profiler-4.0.1/lib/patches/db/pg/alias_method.rb:109:in 'PG::Connection#exec'
/var/www/discourse/vendor/bundle/ruby/3.4.0/gems/rack-mini-profiler-4.0.1/lib/patches/db/pg/alias_method.rb:109:in 'PG::Connection#async_exec'

```

…

```plaintext
from /var/www/discourse/app/models/backup_metadata.rb:16:in 'BackupMetadata.update_last_restore_date'
from /var/www/discourse/lib/backup_restore/database_restorer.rb:31:in 'BackupRestore::DatabaseRestorer#restore'
from /var/www/discourse/lib/backup_restore/restorer.rb:61:in 'BackupRestore::Restorer#run'
from script/discourse:242:in 'DiscourseCLI#restore'

```

…

```plaintext
Trying to rollback...
Cleaning stuff up...
Dropping functions from the discourse_functions schema...
Something went wrong while dropping functions from the discourse_functions schema
PQsocket() can't get socket descriptor

```

### 理论分析

Discourse 使用第二个数据库连接进行实际的恢复操作，并使用另一个连接执行迁移。  
当这些操作完成后，它会在主连接上重新连接到数据库并执行 `BackupMetadata.update_last_restore_date`，但这会立即失败。

导致此失败的原因似乎是数据库重连 实际上并没有重新连接。  
它复用了缓存的 `ConnectionHandler`。参见 [此处](https://github.com/discourse/rails_multisite/blob/86c6ce96e2599d5fe20dd6024f1cdcc76f795da3/lib/rails_multisite/connection_management.rb#L158-L185)。而在 45 分钟后，该连接已经断开。

```ruby
handler = connection_handlers[handler_key(spec)]

unless handler
  handler = ActiveRecord::ConnectionAdapters::ConnectionHandler.new
  handler.establish_connection(spec.config)
  connection_handlers[handler_key(spec)] = handler
end

ActiveRecord::Base.connection_handler = handler

```

### 临时解决方案

Postgres 的 `tcp_keepalives_idle` = 0，这意味着回退到操作系统设置。  
操作系统的 `net.ipv4.tcp_keepalive_time = 7200`（2 小时）。

```plaintext
ALTER SYSTEM SET tcp_keepalives_idle = 60; 
ALTER SYSTEM SET tcp_keepalives_interval = 30; 
ALTER SYSTEM SET tcp_keepalives_count = 5;

```

这可以防止连接被关闭，并解决了该问题。

### 建议的修复方案

让重连代码在重新建立连接之前执行 `ActiveRecord::Base.connection_handler.clear_all_connections!` 或类似操作。

或者，更通用的做法是，为 `establish_connection` 添加一个 `reconnect` 参数，以绕过缓存的 handler。
