我成功修复了这个问题。我特意等了一段时间,确认问题确实已经解决,目前看起来一切正常。
环境配置
Discourse 运行在 Docker 中,并通过 Unix 套接字(/var/discourse/shared/standalone/nginx.http.sock)暴露 nginx。Caddy 作为反向代理位于其前端,并通过 X-Real-IP 传递客户端的真实 IP。
有两处需要修改:容器内的 nginx 配置,以及我运行重建的方式。
1. app.yml
##Plugins go here
## see https://meta.discourse.org/t/19157 for details
hooks:
# This after_code block is just my own plugin list — it has nothing to do with
# the fix. Keep whatever you already have here.
after_code:
- exec:
[...]
# This is the part that matters.
# 1. Writes an http-level map that turns the literal "unix:" into 127.0.0.1.
# The 00- prefix makes nginx load it before discourse.conf.
# 2. Rewrites discourse.conf so X-Forwarded-For uses the mapped variable
# instead of the raw $remote_addr.
# 3. nginx -t fails the build if the result is not valid.
after_web_config:
- exec: >-
printf 'map $remote_addr $safe_remote_addr {\n "unix:" 127.0.0.1;\n default $remote_addr;\n}\nreal_ip_header X-Real-IP;\n'
> /etc/nginx/conf.d/00-safe-remote-addr.conf
- exec: >-
sed -i 's/X-Forwarded-For \$remote_addr;/X-Forwarded-For $safe_remote_addr;/g'
/etc/nginx/conf.d/discourse.conf
- exec: nginx -t
关键代码位于 after_web_config 主体内部。
2. Caddyfile 配置块
下面的脚本通过标志文件(flag file)来驱动维护模式,因此你的站点配置块需要知晓这一点:
forum.example.com {
request_header X-Real-IP {remote_host}
root * /var/caddy/flags
@maintenance file maintenance.flag
handle @maintenance {
respond "Maintenance in progress. We'll be back in a few minutes." 503
}
handle {
reverse_proxy unix//var/discourse/shared/standalone/nginx.http.sock
}
}
脚本中的标志文件路径与 @maintenance 匹配器查找的文件必须是同一个文件。如果你重命名了其中一个,另一个也必须重命名,否则维护模式将静默失效,永远不会触发。
如果同一个 Caddy 实例还服务于其他应用,请使用一个仅由论坛配置块匹配的标志文件,否则 Discourse 的重建会导致其他应用随之宕机。
3. 重建脚本
我不再直接使用 ./launcher rebuild app。我使用的是 discourse-rebuild.sh:
#!/bin/bash
#
# discourse-rebuild.sh — Rebuilds the Discourse container without leaving orphaned
# requests behind and without the `invalid input syntax for type inet: "unix:"` error.
#
# CONTEXT
# Discourse runs in Docker and exposes nginx on a Unix socket
# (/var/discourse/shared/standalone/nginx.http.sock). Caddy acts as a reverse
# proxy in front of it and passes the client's real IP through X-Real-IP.
#
# A `rebuild` destroys the container and recreates the socket with a new inode.
# During that transition there are two problematic windows:
#
# 1) Caddy keeps state from the previous socket until it is reloaded.
# 2) nginx starts accepting connections as soon as it boots, but Unicorn
# takes ~15s longer before it can serve them.
#
# A request landing in either window may arrive with no X-Real-IP. $remote_addr
# is then left holding the literal "unix:", which PostgreSQL rejects when
# inserting it into an inet column -> HTTP 500.
#
# WHAT IT DOES
# 1. Raises a flag file that puts the site into 503 (Caddy checks it on every
# request, so it takes effect instantly and with no reload).
# 2. Waits until nothing but nginx itself is holding the socket open.
# 3. Rebuilds the container.
# 4. Polls /srv/status against the socket until Unicorn answers 200.
# 5. Reloads Caddy so it picks up the new socket, and clears the flag.
# 6. Starts the watcher that logs any leftover "unix:" hit.
#
# If the rebuild fails, or Discourse never answers within the polling window,
# the flag is NOT removed: the site stays in maintenance on purpose, so a broken
# container is never exposed. Bring it back up by hand with:
# rm -f /var/caddy/flags/maintenance.flag
#
# REQUIREMENTS
# - lsof installed, and root privileges.
# - FLAG below must point at the exact same file the Caddyfile @maintenance
# matcher looks for.
# - request_header X-Real-IP {remote_host} in that same Caddyfile block.
#
# USAGE
# ./discourse-rebuild.sh
#
# TO SEE WHAT THE WATCHER CAUGHT
# cat /var/log/unixip-hits.log
#
set -e
FLAG=/var/caddy/flags/maintenance.flag
SOCK=/var/discourse/shared/standalone/nginx.http.sock
cleanup() {
local code=$?
systemctl reload caddy
if [ $code -eq 0 ]; then
rm -f "$FLAG"
echo "✅ Rebuild finished. Site is online."
else
echo "⚠️ Rebuild failed (exit code $code). The site is still in maintenance."
echo " Check it, and once it's ready: rm -f $FLAG"
fi
}
trap cleanup EXIT
mkdir -p "$(dirname "$FLAG")"
touch "$FLAG"
echo "🔧 Maintenance is on. Waiting for in-flight requests..."
# Rough heuristic: count the processes holding the socket open and wait until
# only the listener is left. Up to 30s.
for i in $(seq 30); do
n=$(lsof -t "$SOCK" 2>/dev/null | wc -l || echo 0)
[ "$n" -le 1 ] && break
sleep 1
done
/var/discourse/launcher rebuild app
# Up to 60 attempts: about 2 minutes of sleeps, more if any curl hits its own
# 5s timeout.
echo "⏳ Waiting for Discourse to answer..."
status=000
for i in $(seq 60); do
status=$(curl -s -o /dev/null -w '%{http_code}' --max-time 5 \
--unix-socket "$SOCK" http://localhost/srv/status 2>/dev/null || echo 000)
[ "$status" = "200" ] && break
sleep 2
done
if [ "$status" != "200" ]; then
echo "⚠️ Discourse never answered within the polling window (last status code: $status)."
exit 1
fi
echo "✅ Discourse ready after ~$((i*2))s."
systemd-run --unit=unixip-watch --collect \
/bin/bash -c "docker exec app tail -F /var/log/nginx/access.log | grep --line-buffered 'unix:' >> /var/log/unixip-hits.log"
脚本实际执行的操作
- 通过 Caddy 的标志文件将站点置于维护模式。
- 等待套接字空闲,然后运行
./launcher rebuild app。 - 通过套接字轮询
/srv/status,直到 Unicorn 返回200。这是关键部分:它阻止了请求在 Unicorn 仍在启动时到达 nginx。 - 重新加载 Caddy(双保险,事实证明这一步并非关键)。
- 仅当轮询成功时才清除维护标志。
几周过去,经过多次重建,我再也没有看到过那个错误。
附注:我真的不知道我到底做了什么。