# Restore fails because of non working database reconnect

**URL:** <https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615>\
**Category:** Bug\
**Created:** [September 29, 2026, 11:40pm UTC](https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615 "2026-09-29T23:40:34Z")\
**Posts on this page:** 1\
**Showing post:** 1

<div class="post-metadata">

**Author:** ![RGJ](https://sea3.discourse-cdn.com/meta/user_avatar/meta.discourse.org/rgj/32/523185_2.png) [@RGJ](https://meta.discourse.org/u/RGJ)\
**Post date:** [September 29, 2026, 11:40pm UTC](https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615/1 "2026-09-29T23:40:34Z")

</div>

### Situation:

A two container discourse installation with web and data on different hosts.

### Problem

A large, 22 GB, command line restore from a very old Discourse version (i.e. many lenghty migrations to run) fails after 45 minutes, right after the database restore.

```plaintext
Reconnecting to the database...
EXCEPTION: PQconsumeInput() could not receive data from server: Connection timed out
SSL SYSCALL error: Connection timed out
/var/www/discourse/vendor/bundle/ruby/3.4.0/gems/rack-mini-profiler-4.0.1/lib/patches/db/pg/alias_method.rb:109:in 'PG::Connection#exec'
/var/www/discourse/vendor/bundle/ruby/3.4.0/gems/rack-mini-profiler-4.0.1/lib/patches/db/pg/alias_method.rb:109:in 'PG::Connection#async_exec'

```

…

```plaintext
from /var/www/discourse/app/models/backup_metadata.rb:16:in 'BackupMetadata.update_last_restore_date'
from /var/www/discourse/lib/backup_restore/database_restorer.rb:31:in 'BackupRestore::DatabaseRestorer#restore'
from /var/www/discourse/lib/backup_restore/restorer.rb:61:in 'BackupRestore::Restorer#run'
from script/discourse:242:in 'DiscourseCLI#restore'

```

…

```plaintext
Trying to rollback...
Cleaning stuff up...
Dropping functions from the discourse_functions schema...
Something went wrong while dropping functions from the discourse_functions schema
PQsocket() can't get socket descriptor

```

### Theory

Discourse uses a second database connection for the actual restore and another one for the migration.  
When those are finished, it reconnects to the database on its primary connection and performs `BackupMetadata.update_last_restore_date` which immediately fails.

The reason for this failure seems to be that the database reconnect does not actually reconnect.  
It reuses the cached `ConnectionHandler`. See [here](https://github.com/discourse/rails_multisite/blob/86c6ce96e2599d5fe20dd6024f1cdcc76f795da3/lib/rails_multisite/connection_management.rb#L158-L185). And that connection is gone after 45 minutes.

```ruby
handler = connection_handlers[handler_key(spec)]

unless handler
  handler = ActiveRecord::ConnectionAdapters::ConnectionHandler.new
  handler.establish_connection(spec.config)
  connection_handlers[handler_key(spec)] = handler
end

ActiveRecord::Base.connection_handler = handler

```

### Workaround

Postgres has `tcp_keepalives_idle` = 0 which means fall back to the OS setting.  
OS has `net.ipv4.tcp_keepalive_time = 7200` (2 hours)

```plaintext
ALTER SYSTEM SET tcp_keepalives_idle = 60; 
ALTER SYSTEM SET tcp_keepalives_interval = 30; 
ALTER SYSTEM SET tcp_keepalives_count = 5;

```

Keeps the connection from getting closed and resolves the issue.

### Suggested fix

Have the reconnection code do

`ActiveRecord::Base.connection_handler.clear_all_connections!` or similar before re-establishing the connection.

Or, more generic, add a `reconnect` parameter to `establish_connection` which bypasses the cached handler.

---

_[View the full topic](https://meta.discourse.org/t/restore-fails-because-of-non-working-database-reconnect/413615)._
