CaddyとDiscourseの間にTCPではなくUnixソケットを使用していますが、時々「unix:」エラーが発生します

ログには以下が表示されています:

PG::InvalidTextRepresentation (ERROR: invalid input syntax for type inet: "unix:" LINE 7: client_ip = 'unix:', ^ ) lib/mini_sql_multisite_connection.rb:109:in 'MiniSqlMult

Job exception: ERROR: invalid input syntax for type inet: "unix:" LINE 2: SET ip_address = 'unix:' ^  

発生すると、**「Oops - Error 500」**ページが表示されます。最初は、Caddy がクライアントの IP アドレスを Discourse に転送していなかったため、Caddy が原因ではないかと思いました。そのため、いくつかの異なる設定を試しましたが、いずれも問題を解決できませんでした。

なお、この問題は頻繁に発生するわけではありません。たまに発生するだけで、発生した際はページを単純にリロードするだけでサイトはすぐに回復します。その後、しばらく正常に動作し、再びランダムにエラーが発生します。

私は専門家ではありませんし、セルフホストのインスタンスでは Caddy ではなく Nginx を使用していますが、./app/containers.yml の実際の Caddy と Discourse テンプレートの設定を教えていただけますか?

エラーが発生する前に、普段どのような操作を行っていたか覚えていますか?例えば、管理者設定の変更、投稿、プラグインの使用などです。

あなたの問題に具体的な解決策をお伝えできないことをお詫びしますが、あなたの返信がコミュニティから明確で迅速な回答を得るために、質問に価値を加えるものになると考えています。

これは、リクエスト/ユーザーのリモートIPがあなたのソケットを指していることを意味し、プロキシチェーンのどこかが誤って設定されていることを示しています。

@satonotdead @Falco

app.yml

templates:
  - "templates/postgres.18.template.yml"
  - "templates/redis.template.yml"
  - "templates/web.template.yml"
  
  - "templates/web.socketed.template.yml"
  
  - "templates/enable-ruby-yjit.yml"


## which TCP/IP ports should this container expose?
## If you want Discourse to share a port with another webserver like Apache or nginx,
## see https://meta.discourse.org/t/17247 for details
expose:
  # - "66:80"   # http
  # - "66:443" # https

params:
  ## Which Git revision should this container use? (default: latest)
  version: latest
  ## Maximum upload size (default: 10m)
  upload_size: 150m
  
  db_default_text_search_config: "pg_catalog.english"

  ## Set db_shared_buffers to a max of 25% of the total memory.
  ## will be set automatically by bootstrap based on detected RAM, or you can override
  db_shared_buffers: "2048MB"

  ## can improve sorting performance, but adds memory usage per-connection
  #db_work_mem: "40MB"

  ## Which Git revision should this container use? (default: tests-passed)
  #version: tests-passed

env:
  LC_ALL: en_US.UTF-8
  LANG: en_US.UTF-8
  LANGUAGE: en_US.UTF-8
  # DISCOURSE_DEFAULT_LOCALE: en

  ## https://meta.discourse.org/t/rescaling-the-server-which-configs-need-to-be-changed-unicorn-workers-memory-etc/252788
  ## How many concurrent web requests are supported? Depends on memory and CPU cores.
  ## will be set automatically by bootstrap based on detected CPUs, or you can override
  UNICORN_WORKERS: 8

  ## TODO: The domain name this Discourse instance will respond to
  ## Required. Discourse will not work with a bare IP number.
  DISCOURSE_HOSTNAME: example.com

  ## Uncomment if you want the container to be started with the same
  ## hostname (-h option) as specified above (default "$hostname-$config")
  #DOCKER_USE_HOSTNAME: true

  ## TODO: List of comma delimited emails that will be made admin and developer
  ## on initial signup example 'user1@example.com,user2@example.com'
  DISCOURSE_DEVELOPER_EMAILS: 'admin+discourse@example.com'

  ## TODO: The SMTP mail server used to validate new accounts and send notifications
  # SMTP ADDRESS is required
  # WARNING: SMTP password should be wrapped in quotes to avoid problems
  DISCOURSE_SMTP_ADDRESS: smtp.provider.com
  DISCOURSE_SMTP_PORT: 587
  DISCOURSE_SMTP_USER_NAME: noreply@example.com
  DISCOURSE_SMTP_PASSWORD: "***"
  #DISCOURSE_SMTP_ENABLE_START_TLS: true           # (optional, default: true)
  DISCOURSE_SMTP_DOMAIN: example.com # (required by some providers)
  DISCOURSE_NOTIFICATION_EMAIL: noreply@example.com
  #DISCOURSE_SMTP_OPENSSL_VERIFY_MODE: peer        # (optional, default: peer, valid values: none, peer, client_once, fail_if_no_peer_cert)
  #DISCOURSE_SMTP_AUTHENTICATION: plain            # (default: plain, valid values: plain, login, cram_md5)

  ## If you added the Lets Encrypt template, uncomment below to get a free SSL certificate
  # LETSENCRYPT_ACCOUNT_EMAIL: admin+letsencrypt@example.com

  ## The http or https CDN address for this Discourse instance (configured to pull)
  ## see https://meta.discourse.org/t/14857 for details
  #DISCOURSE_CDN_URL: https://discourse-cdn.example.com

  ## The maxmind geolocation IP account ID and license key for IP address lookups
  ## see https://meta.discourse.org/t/-/173941 for details
  #DISCOURSE_MAXMIND_ACCOUNT_ID: 123456
  #DISCOURSE_MAXMIND_LICENSE_KEY: 1234567890123456

  # Force HTTPS
  DISCOURSE_FORCE_HTTPS: true
  
  # Request Limits
  DISCOURSE_MAX_REQS_PER_IP_MODE: none
  DISCOURSE_MAX_ADMIN_API_REQS_PER_MINUTE: 12000
  
  DISCOURSE_MAX_DATA_EXPLORER_API_REQ_MODE: none
  DISCOURSE_MAX_DATA_EXPLORER_API_REQS_PER_10_SECONDS: 1000

  DISCOURSE_YJIT_ENABLED: true

## The Docker container is stateless; all data is stored in /shared
volumes:
  - volume:
      host: /var/discourse/shared/standalone
      guest: /shared
  - volume:
      host: /var/discourse/shared/standalone/log/var-log
      guest: /var/log
  - volume:
      host: /var/discourse/plugins
      guest: /var/plugins

## Plugins go here
## see https://meta.discourse.org/t/19157 for details
hooks:
  after_code:
    - exec:
        cd: $home/plugins
        cmd:
          - git clone https://github.com/discourse/docker_manager.git
          - cp -a /var/plugins/. $home/plugins/

## Any custom commands to run after building
run:
  - exec: echo "Beginning of custom commands"
  ## If you want to set the 'From' email address for your first registration, uncomment and change:
  ## After getting the first signup email, re-comment the line. It only needs to run once.
  #- exec: rails r "SiteSetting.notification_email='info@unconfigured.discourse.org'"
  - exec: echo "End of custom commands"

Caddyfile

#---------------------------------------------------------------------------------

{
	storage file_system {
		root /var/caddy/data
	}
}

#---------------------------------------------------------------------------------

# To achieve an A+ rating in SSL security
(hsts) {
	header {
		Strict-Transport-Security "max-age=31536000; includeSubDomains"
	}
}

# self-hosted acme-dns credentials
(tls-challenge) {
	tls admin+letsencrypt@example.com {
		dns acmedns {
			username ***
			password ***
			subdomain ***
			server_url http://acme.example.com:5005
		}
	}
}

#---------------------------------------------------------------------------------
# Forum (Discourse)
#---------------------------------------------------------------------------------
example.com {
	#-------------------------------------------------------------------------------
	import hsts
	import tls-challenge
	#-------------------------------------------------------------------------------
	request_header X-Real-IP {remote_host}
	reverse_proxy unix//var/discourse/shared/standalone/nginx.http.sock
	#-------------------------------------------------------------------------------
}

#---------------------------------------------------------------------------------
# Subdominios
#---------------------------------------------------------------------------------

#---------------------------------------------------------------------------------
*.example.com {
	#-------------------------------------------------------------------------------
	import hsts
	import tls-challenge
	#-------------------------------------------------------------------------------

	# Force the non-www version of the domain
	@www host www.example.com
	redir @www https://example.com{uri} permanent
  
	#-------------------------------------------------------------------------------
  
	@store host store.example.com
	handle @store {
		@wc_private {
			path /license.txt
			path /readme.html

			# Archivos críticos
			path /wp-config.php
			path /xmlrpc.php
			path /wp-settings.php
			path /wp-load.php
			path /wp-blog-header.php

			path /wp-admin.php

			path /wp-admin/install.php

			path /wp-content/uploads/wc-logs/*
			path /wp-content/uploads/nuvei-logs/*

			path /wp-content/uploads/woocommerce_uploads/*
		}

		# protect_php
		@wc_private_php {
			not path /wp-includes/ms-files.php

			path_regexp protect_php ^/(wp-includes|wp-admin/includes|wp-content/uploads)/.*\.php$
		}
		respond @wc_private 403
		respond @wc_private_php 403

		root * /usr/share/wordpress
		php_fastcgi unix//run/php/php8.3-fpm.sock
		file_server
	}
  
  #-------------------------------------------------------------------------------
  # Multiple standalone HTML files
  #-------------------------------------------------------------------------------

  @join host join.example.com
  handle @join {
    root * /var/example/join
    file_server
  }

  #-------------------------------------------------------------------------------
  # No matchers
  #-------------------------------------------------------------------------------

	handle {
		respond 404
	}

	#-------------------------------------------------------------------------------
}

より具体的には、このエラーは前回のエラーから数時間後に発生し、頻繁に発生するわけではありません。app.ymlやCaddyfileに異常は見当たりません。

ソケットへの接続が切断されるまで正常であれば、明示的なタイムアウトとトランスポートオプションを追加できると思います。

明示的なタイムアウトとトランスポートオプションを追加できると思います

どうすればよいですか?

さて、以前にも述べましたが、私は Caddy を使用していませんが、公式ドキュメントを参照することができます。

問題を発見したと思います。リビルドを実行しても、Caddy を再起動しません。今後は以下のようにします:

./launcher rebuild app && systemctl reload caddy

Caddy はリビルド後にソケットキャッシュや同様のものを使用しているのではないかと推測しています。ただし、通常リビルド直後に発生することはありません。様子を見てみましょう。

何日もの間、これに取り組んできました。Caddyがヘッダーを送信し、nginxがそれを受信して適用していることを確認しました。CDNも、ソケットに干渉する他のプロセスも、Webhookも存在しません。すべてのテストがパスしています。しかし、Unixのエラーが引き続き発生します。ビルドの再構築が原因だと思ったのですが、そうではありません。これらすべては、非常に長い期間にわたってランダムに発生しています。

今、TCPが唯一の選択肢なのでしょうか?

私が提案したタイムアウトを Caddy のテンプレートに追加しましたか?私は専門家ではありませんが、それならあなたの問題は解決すると思います。

そのような做法は意味がありません。一部のリクエストは15秒で失敗し、他のリクエストは1秒未満で失敗するためです。基本的には無意味です。

はい、以下の2つのリンクを確認できます:

ヘッダーが不足しているか、またはDiscourseテンプレートファイルへのチェーンが不足しているようです。

@Falco @satonotdead

なんとか修正できました。本当に直ったかどうか確認するためにしばらく様子を見てみましたが、おそらく直ったようです。

セットアップ

Discourse は Docker 上で動作し、nginx を Unix ソケット (/var/discourse/shared/standalone/nginx.http.sock) で公開しています。Caddy がその前面にリバースプロキシとして配置され、X-Real-IP ヘッダーを通じてクライアントの実 IP を転送します。

変更が必要だったのは2点です:コンテナ内の nginx 設定と、リビルドを実行する方法です。

1. app.yml

## Plugins go here
## see https://meta.discourse.org/t/19157 for details
hooks:
  # This after_code block is just my own plugin list — it has nothing to do with
  # the fix. Keep whatever you already have here.
  after_code:
    - exec:
        [...]

  # This is the part that matters.
  #   1. Writes an http-level map that turns the literal "unix:" into 127.0.0.1.
  #      The 00- prefix makes nginx load it before discourse.conf.
  #   2. Rewrites discourse.conf so X-Forwarded-For uses the mapped variable
  #      instead of the raw $remote_addr.
  #   3. nginx -t fails the build if the result is not valid.
  after_web_config:
    - exec: >-
        printf 'map $remote_addr $safe_remote_addr {\n  "unix:" 127.0.0.1;\n  default $remote_addr;\n}\nreal_ip_header X-Real-IP;\n'
        > /etc/nginx/conf.d/00-safe-remote-addr.conf
    - exec: >-
        sed -i 's/X-Forwarded-For \$remote_addr;/X-Forwarded-For $safe_remote_addr;/g'
        /etc/nginx/conf.d/discourse.conf
    - exec: nginx -t

重要なコードは、after_web_config の本体の中にあります。

2. Caddyfile のブロック

以下のスクリプトはフラグファイルを通じてメンテナンスモードを制御するため、サイトブロックがそれを知る必要があります:

forum.example.com {
    request_header X-Real-IP {remote_host}

    root * /var/caddy/flags
    @maintenance file maintenance.flag
    handle @maintenance {
        respond "Maintenance in progress. We'll be back in a few minutes." 503
    }

    handle {
        reverse_proxy unix//var/discourse/shared/standalone/nginx.http.sock
    }
}

スクリプト内のフラグパスと、@maintenance マッチャーが参照するファイルは同じファイルである必要があります。片方をリネームしたら、もう片方もリネームしてください。そうしないと、メンテナンスモードがサイレントに発動しなくなります。

同じ Caddy インスタンスが他のアプリも配信している場合、フォーラムのブロックのみがマッチするフラグを使用してください。そうしないと、Discourse のリビルド時に他のアプリも一緒にダウンしてしまいます。

3. リビルドスクリプト

もう ./launcher rebuild app を直接使いません。代わりに discourse-rebuild.sh を使用しています:

#!/bin/bash
#
# discourse-rebuild.sh — Rebuilds the Discourse container without leaving orphaned
# requests behind and without the `invalid input syntax for type inet: "unix:"` error.
#
# CONTEXT
#   Discourse runs in Docker and exposes nginx on a Unix socket
#   (/var/discourse/shared/standalone/nginx.http.sock). Caddy acts as a reverse
#   proxy in front of it and passes the client's real IP through X-Real-IP.
#
#   A `rebuild` destroys the container and recreates the socket with a new inode.
#   During that transition there are two problematic windows:
#
#     1) Caddy keeps state from the previous socket until it is reloaded.
#     2) nginx starts accepting connections as soon as it boots, but Unicorn
#        takes ~15s longer before it can serve them.
#
#   A request landing in either window may arrive with no X-Real-IP. $remote_addr
#   is then left holding the literal "unix:", which PostgreSQL rejects when
#   inserting it into an inet column -> HTTP 500.
#
# WHAT IT DOES
#   1. Raises a flag file that puts the site into 503 (Caddy checks it on every
#      request, so it takes effect instantly and with no reload).
#   2. Waits until nothing but nginx itself is holding the socket open.
#   3. Rebuilds the container.
#   4. Polls /srv/status against the socket until Unicorn answers 200.
#   5. Reloads Caddy so it picks up the new socket, and clears the flag.
#   6. Starts the watcher that logs any leftover "unix:" hit.
#
#   If the rebuild fails, or Discourse never answers within the polling window,
#   the flag is NOT removed: the site stays in maintenance on purpose, so a broken
#   container is never exposed. Bring it back up by hand with:
#       rm -f /var/caddy/flags/maintenance.flag
#
# REQUIREMENTS
#   - lsof installed, and root privileges.
#   - FLAG below must point at the exact same file the Caddyfile @maintenance
#     matcher looks for.
#   - request_header X-Real-IP {remote_host} in that same Caddyfile block.
#
# USAGE
#   ./discourse-rebuild.sh
#
# TO SEE WHAT THE WATCHER CAUGHT
#   cat /var/log/unixip-hits.log
#

set -e

FLAG=/var/caddy/flags/maintenance.flag
SOCK=/var/discourse/shared/standalone/nginx.http.sock

cleanup() {
  local code=$?
  systemctl reload caddy
  if [ $code -eq 0 ]; then
    rm -f "$FLAG"
    echo "✅ Rebuild finished. Site is online."
  else
    echo "⚠️  Rebuild failed (exit code $code). The site is still in maintenance."
    echo "    Check it, and once it's ready: rm -f $FLAG"
  fi
}
trap cleanup EXIT

mkdir -p "$(dirname "$FLAG")"
touch "$FLAG"
echo "🔧 Maintenance is on. Waiting for in-flight requests..."

# Rough heuristic: count the processes holding the socket open and wait until
# only the listener is left. Up to 30s.
for i in $(seq 30); do
  n=$(lsof -t "$SOCK" 2>/dev/null | wc -l || echo 0)
  [ "$n" -le 1 ] && break
  sleep 1
done

/var/discourse/launcher rebuild app

# Up to 60 attempts: about 2 minutes of sleeps, more if any curl hits its own
# 5s timeout.
echo "⏳ Waiting for Discourse to answer..."
status=000
for i in $(seq 60); do
  status=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 \
    --unix-socket "$SOCK" http://localhost/srv/status 2>/dev/null || echo 000)
  [ "$status" = "200" ] && break
  sleep 2
done

if [ "$status" != "200" ]; then
  echo "⚠️  Discourse never answered within the polling window (last status code: $status)."
  exit 1
fi

echo "✅ Discourse ready after ~$((i*2))s."

systemd-run --unit=unixip-watch --collect \
  /bin/bash -c "docker exec app tail -F /var/log/nginx/access.log | grep --line-buffered 'unix:' >> /var/log/unixip-hits.log"

スクリプトが実際にやっていること

  1. Caddy のフラグファイルを通じてサイトをメンテナンスモードにします。
  2. ソケットが静まるのを待ってから ./launcher rebuild app を実行します。
  3. Unicorn が 200 を返すまでソケット経由で /srv/status をポーリングします。これが重要な部分です:Unicorn がまだ起動中の間にリクエストが nginx に到達することを防ぐためです。
  4. Caddy をリロードします(念のため。実際には重要ではなかった)。
  5. そのポーリングが成功した場合のみ、メンテナンスフラグをクリアします。

数週間、数回のリビルドを経て、そのエラーはもう見かけません。

P.S. 一体何をしたのか、自分でもまったくわかりません。